{
  "id": 342992,
  "title": "Is 0.801 a bottleneck?  ",
  "url": "/competitions/amex-default-prediction/discussion/342992",
  "author_name": "zhehao liang",
  "post_date": "2022-08-09T14:19:04.123000",
  "votes": 9,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Hi. Can someone technically explain why all model comes to 0.801 and (or very closely)?<br>\nIs that common in other competition?</p>",
  "messages": [
    {
      "id": 1891620,
      "postDate": "2022-08-09T14:19:04.123Z",
      "content": "<p>Hi. Can someone technically explain why all model comes to 0.801 and (or very closely)?<br>\nIs that common in other competition?</p>",
      "rawMarkdown": "Hi. Can someone technically explain why all model comes to 0.801 and (or very closely)?\nIs that common in other competition?",
      "votes": 9
    },
    {
      "id": 1891837,
      "postDate": "2022-08-09T17:20:34.397Z",
      "content": "<p>It is not common in all competitions, but it happens when the best public model is so close to the best model overall. I think as a group we could have gotten a slightly better score had noise not been introduced into data, and if we knew the exact meaning of all features. But chances are very good that for this particular metrics, we already have a score that is as good as it gets. If you think about it, a huge amount of time and computer resources have been spent (and likely wasted) on improving a score on the third decimal place, which in real world would never happen. Prizes and medals are big incentives.</p>",
      "rawMarkdown": "It is not common in all competitions, but it happens when the best public model is so close to the best model overall. I think as a group we could have gotten a slightly better score had noise not been introduced into data, and if we knew the exact meaning of all features. But chances are very good that for this particular metrics, we already have a score that is as good as it gets. If you think about it, a huge amount of time and computer resources have been spent (and likely wasted) on improving a score on the third decimal place, which in real world would never happen. Prizes and medals are big incentives.",
      "votes": 8
    },
    {
      "id": 1893459,
      "postDate": "2022-08-10T20:40:37.360Z",
      "content": "<p>In this competetion, the cv score and the LB score basically correspond. In my experiments, this correspondence becomes less reliable after my cv is higher than 0.8005 when LB reached 0.801. I think LB 0.801 might not the bottleneck, but our models are stucked in local optima. But time is running out and hopefully someone can break 0.801.</p>",
      "rawMarkdown": "In this competetion, the cv score and the LB score basically correspond. In my experiments, this correspondence becomes less reliable after my cv is higher than 0.8005 when LB reached 0.801. I think LB 0.801 might not the bottleneck, but our models are stucked in local optima. But time is running out and hopefully someone can break 0.801.",
      "votes": 4,
      "replies": [
        {
          "id": 1896393,
          "postDate": "2022-08-12T19:36:00.617Z",
          "content": "<p>I think it's because of the non-uniform number of monthly statements for users. Acc to my experiments, models are performing well in 0.82+ range for the users where we have 13 records for them. However it significantly drops for users with less than 13 records for them. </p>\n<p>This is how my CV scores look for different set of users having varied number of training records.</p>\n<p><code>Nunique 1: less-than score 0.6225173217635962</code><br>\n<code>Nunique 2: less-than score 0.6226079728530806</code><br>\n<code>Nunique 3: less-than score 0.6346767585761496</code><br>\n<code>Nunique 4: less-than score 0.6300117792947451</code><br>\n<code>Nunique 5: less-than score 0.6365798250458007</code><br>\n<code>Nunique 6: less-than score 0.6426331662786908</code><br>\n<code>Nunique 7: less-than score 0.6485652488655076</code><br>\n<code>Nunique 8: less-than score 0.6518806963951781</code><br>\n<code>Nunique 9: less-than score 0.6570058871126497</code><br>\n<code>Nunique 10: less-than score 0.659927197311278</code><br>\n<code>Nunique 11: less-than score 0.6631892002647499</code><br>\n<code>Nunique 12: less-than score 0.6719086819329969</code><br>\n<code>Nunique 13: less-than score 0.7975376272588612</code></p>\n<p>After checking the user distribution between public and private, I think private scores will be more than the public ones as the percentage of users having 13 records is more in private as compared to public.</p>",
          "rawMarkdown": "I think it's because of the non-uniform number of monthly statements for users. Acc to my experiments, models are performing well in 0.82+ range for the users where we have 13 records for them. However it significantly drops for users with less than 13 records for them. \n\nThis is how my CV scores look for different set of users having varied number of training records.\n\n```Nunique 1: less-than score 0.6225173217635962```\n```Nunique 2: less-than score 0.6226079728530806```\n```Nunique 3: less-than score 0.6346767585761496```\n```Nunique 4: less-than score 0.6300117792947451```\n```Nunique 5: less-than score 0.6365798250458007```\n```Nunique 6: less-than score 0.6426331662786908```\n```Nunique 7: less-than score 0.6485652488655076```\n```Nunique 8: less-than score 0.6518806963951781```\n```Nunique 9: less-than score 0.6570058871126497```\n```Nunique 10: less-than score 0.659927197311278```\n```Nunique 11: less-than score 0.6631892002647499```\n```Nunique 12: less-than score 0.6719086819329969```\n```Nunique 13: less-than score 0.7975376272588612```\n\nAfter checking the user distribution between public and private, I think private scores will be more than the public ones as the percentage of users having 13 records is more in private as compared to public.",
          "votes": 15
        },
        {
          "id": 1904010,
          "postDate": "2022-08-17T21:18:43.347Z",
          "content": "<p>Thanks, that's a good explanation why public LB's score is slightly higher than CV.</p>",
          "rawMarkdown": "Thanks, that's a good explanation why public LB's score is slightly higher than CV.",
          "votes": 2
        },
        {
          "id": 1904865,
          "postDate": "2022-08-18T14:54:06.160Z",
          "content": "<p>Hi, just a question is oversampling the data for customers with less than 13 rows a viable option to to increase the score? For example removing the first transaction from a customer with 13 rows and adding it to a list of customers with 12 rows, removing 2 transactions from 13 and removing 1 transaction from 12 and adding it to customers with 11 rows. Could it also be possible to use the oversampling to do something like GroupKFold or Stratified KFold based on the customer rows? I just wanna because I wanted to experiment on this but I am begginner and kaggle notebooks, just keep throwing cuda allocation error when I try.</p>",
          "rawMarkdown": "Hi, just a question is oversampling the data for customers with less than 13 rows a viable option to to increase the score? For example removing the first transaction from a customer with 13 rows and adding it to a list of customers with 12 rows, removing 2 transactions from 13 and removing 1 transaction from 12 and adding it to customers with 11 rows. Could it also be possible to use the oversampling to do something like GroupKFold or Stratified KFold based on the customer rows? I just wanna because I wanted to experiment on this but I am begginner and kaggle notebooks, just keep throwing cuda allocation error when I try."
        },
        {
          "id": 1904918,
          "postDate": "2022-08-18T15:36:05.533Z",
          "content": "<p>Yes you can try that approach. I did try but didn't work for me. got CV 0.79425 and LB 0.795</p>",
          "rawMarkdown": "Yes you can try that approach. I did try but didn't work for me. got CV 0.79425 and LB 0.795",
          "votes": 1
        },
        {
          "id": 1904945,
          "postDate": "2022-08-18T16:02:49.967Z",
          "content": "<p>Anyway thanks, I dont know if this helps but amex currently uses GRU and GBDT ensemble, in a video I found.</p>",
          "rawMarkdown": "Anyway thanks, I dont know if this helps but amex currently uses GRU and GBDT ensemble, in a video I found."
        }
      ]
    },
    {
      "id": 1891831,
      "postDate": "2022-08-09T17:18:34.047Z",
      "content": "<p>Though I think 0.801 is my limit, I feel like 0.802 is possible, considering that there are so many things that can be done…</p>",
      "rawMarkdown": "Though I think 0.801 is my limit, I feel like 0.802 is possible, considering that there are so many things that can be done...",
      "votes": 4
    },
    {
      "id": 1912599,
      "postDate": "2022-08-24T21:12:07.990Z",
      "content": "<p>Anonymized data necessarily reduces the accuracy of the predictions. Look at this notebook's explanation about B_19 and S_13. <a href=\"https://www.kaggle.com/code/ambrosm/amex-eda-which-makes-sense/notebook\" target=\"_blank\">https://www.kaggle.com/code/ambrosm/amex-eda-which-makes-sense/notebook</a>  (below cell 17) </p>\n<p>The <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> point about diminishing returns seems poignant. Was happy to see someone break .802</p>",
      "rawMarkdown": "Anonymized data necessarily reduces the accuracy of the predictions. Look at this notebook's explanation about B_19 and S_13. https://www.kaggle.com/code/ambrosm/amex-eda-which-makes-sense/notebook  (below cell 17) \n\nThe @roberthatch point about diminishing returns seems poignant. Was happy to see someone break .802",
      "votes": 1
    },
    {
      "id": 1904506,
      "postDate": "2022-08-18T08:28:26.097Z",
      "content": "<p>As of now, public lb first place has been achieved .802</p>",
      "rawMarkdown": "As of now, public lb first place has been achieved .802",
      "votes": 1,
      "replies": [
        {
          "id": 1904838,
          "postDate": "2022-08-18T14:35:14.723Z",
          "content": "<p>Yah, I also achieved .799 today, sorry for being late.</p>",
          "rawMarkdown": "Yah, I also achieved .799 today, sorry for being late.",
          "votes": 1
        },
        {
          "id": 1904858,
          "postDate": "2022-08-18T14:47:22.160Z",
          "content": "<p>It's okay my friend. We live in different time zones and as long as we're still on the forum, we're never late.Good luck to you!</p>",
          "rawMarkdown": "It's okay my friend. We live in different time zones and as long as we're still on the forum, we're never late.Good luck to you!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1892129,
      "postDate": "2022-08-09T21:39:59.957Z",
      "content": "<p>It's also possible we would've noticed gradual continual improvement if 4 digits were being reported. </p>\n<p>Without the fourth digit it's impossible to speculate how possible 802 is. Maybe the top teams would have a decent guess based on their own CV scores, but most of us have zero info to go on. The top score might be. 8014, but could also be .80199. </p>\n<p>We can guess 803 is looking super unlikely though ;). Diminishing returns means that would be a long long long ways past 802. </p>",
      "rawMarkdown": "It's also possible we would've noticed gradual continual improvement if 4 digits were being reported. \n\nWithout the fourth digit it's impossible to speculate how possible 802 is. Maybe the top teams would have a decent guess based on their own CV scores, but most of us have zero info to go on. The top score might be. 8014, but could also be .80199. \n\nWe can guess 803 is looking super unlikely though ;). Diminishing returns means that would be a long long long ways past 802. ",
      "votes": 1
    },
    {
      "id": 1891813,
      "postDate": "2022-08-09T16:57:23.763Z",
      "content": "<p>This seems like a bottleneck! Hope someone breaks it!</p>",
      "rawMarkdown": "This seems like a bottleneck! Hope someone breaks it!",
      "votes": 1
    },
    {
      "id": 1904083,
      "postDate": "2022-08-18T00:09:47.843Z",
      "content": "<p>Obviously not :P</p>",
      "rawMarkdown": "Obviously not :P",
      "votes": 2
    },
    {
      "id": 1900527,
      "postDate": "2022-08-16T05:04:20.617Z",
      "content": "<p>This is dream!</p>",
      "rawMarkdown": "This is dream!",
      "votes": 2
    },
    {
      "id": 1900385,
      "postDate": "2022-08-16T01:45:43.377Z",
      "content": "<p>I think it really depends on the underlying probability distribution. </p>\n<p>let's imagine it was a simple bimodal and either it was a 5% chance of default or 95% </p>\n<p>then the metric maxes out at 80-85, just because there are a bunch of positives in the data from people with a 5% chance of default, your model would rate them low as it should but it reduces the overall possible score</p>",
      "rawMarkdown": "I think it really depends on the underlying probability distribution. \n\nlet's imagine it was a simple bimodal and either it was a 5% chance of default or 95% \n\nthen the metric maxes out at 80-85, just because there are a bunch of positives in the data from people with a 5% chance of default, your model would rate them low as it should but it reduces the overall possible score\n\n\n\n"
    },
    {
      "id": 1894077,
      "postDate": "2022-08-11T08:55:06.407Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1891837,
      "author_name": "Tilii",
      "author_url": "",
      "post_date": "2022-08-09T17:20:34.397000",
      "content": "<p>It is not common in all competitions, but it happens when the best public model is so close to the best model overall. I think as a group we could have gotten a slightly better score had noise not been introduced into data, and if we knew the exact meaning of all features. But chances are very good that for this particular metrics, we already have a score that is as good as it gets. If you think about it, a huge amount of time and computer resources have been spent (and likely wasted) on improving a score on the third decimal place, which in real world would never happen. Prizes and medals are big incentives.</p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 1893459,
      "author_name": "Mengfei Li",
      "author_url": "",
      "post_date": "2022-08-10T20:40:37.360000",
      "content": "<p>In this competetion, the cv score and the LB score basically correspond. In my experiments, this correspondence becomes less reliable after my cv is higher than 0.8005 when LB reached 0.801. I think LB 0.801 might not the bottleneck, but our models are stucked in local optima. But time is running out and hopefully someone can break 0.801.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1896393,
          "author_name": "tarick.morty",
          "author_url": "",
          "post_date": "2022-08-12T19:36:00.617000",
          "content": "<p>I think it's because of the non-uniform number of monthly statements for users. Acc to my experiments, models are performing well in 0.82+ range for the users where we have 13 records for them. However it significantly drops for users with less than 13 records for them. </p>\n<p>This is how my CV scores look for different set of users having varied number of training records.</p>\n<p><code>Nunique 1: less-than score 0.6225173217635962</code><br>\n<code>Nunique 2: less-than score 0.6226079728530806</code><br>\n<code>Nunique 3: less-than score 0.6346767585761496</code><br>\n<code>Nunique 4: less-than score 0.6300117792947451</code><br>\n<code>Nunique 5: less-than score 0.6365798250458007</code><br>\n<code>Nunique 6: less-than score 0.6426331662786908</code><br>\n<code>Nunique 7: less-than score 0.6485652488655076</code><br>\n<code>Nunique 8: less-than score 0.6518806963951781</code><br>\n<code>Nunique 9: less-than score 0.6570058871126497</code><br>\n<code>Nunique 10: less-than score 0.659927197311278</code><br>\n<code>Nunique 11: less-than score 0.6631892002647499</code><br>\n<code>Nunique 12: less-than score 0.6719086819329969</code><br>\n<code>Nunique 13: less-than score 0.7975376272588612</code></p>\n<p>After checking the user distribution between public and private, I think private scores will be more than the public ones as the percentage of users having 13 records is more in private as compared to public.</p>",
          "votes": 15,
          "replies": []
        },
        {
          "id": 1904010,
          "author_name": "Mengfei Li",
          "author_url": "",
          "post_date": "2022-08-17T21:18:43.347000",
          "content": "<p>Thanks, that's a good explanation why public LB's score is slightly higher than CV.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1904865,
          "author_name": "Tarrasque9",
          "author_url": "",
          "post_date": "2022-08-18T14:54:06.160000",
          "content": "<p>Hi, just a question is oversampling the data for customers with less than 13 rows a viable option to to increase the score? For example removing the first transaction from a customer with 13 rows and adding it to a list of customers with 12 rows, removing 2 transactions from 13 and removing 1 transaction from 12 and adding it to customers with 11 rows. Could it also be possible to use the oversampling to do something like GroupKFold or Stratified KFold based on the customer rows? I just wanna because I wanted to experiment on this but I am begginner and kaggle notebooks, just keep throwing cuda allocation error when I try.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1904918,
          "author_name": "tarick.morty",
          "author_url": "",
          "post_date": "2022-08-18T15:36:05.533000",
          "content": "<p>Yes you can try that approach. I did try but didn't work for me. got CV 0.79425 and LB 0.795</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1904945,
          "author_name": "Tarrasque9",
          "author_url": "",
          "post_date": "2022-08-18T16:02:49.967000",
          "content": "<p>Anyway thanks, I dont know if this helps but amex currently uses GRU and GBDT ensemble, in a video I found.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1891831,
      "author_name": "Fritz Cremer",
      "author_url": "",
      "post_date": "2022-08-09T17:18:34.047000",
      "content": "<p>Though I think 0.801 is my limit, I feel like 0.802 is possible, considering that there are so many things that can be done…</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1912599,
      "author_name": "Todd Gardiner",
      "author_url": "",
      "post_date": "2022-08-24T21:12:07.990000",
      "content": "<p>Anonymized data necessarily reduces the accuracy of the predictions. Look at this notebook's explanation about B_19 and S_13. <a href=\"https://www.kaggle.com/code/ambrosm/amex-eda-which-makes-sense/notebook\" target=\"_blank\">https://www.kaggle.com/code/ambrosm/amex-eda-which-makes-sense/notebook</a>  (below cell 17) </p>\n<p>The <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> point about diminishing returns seems poignant. Was happy to see someone break .802</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1904506,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-18T08:28:26.097000",
      "content": "<p>As of now, public lb first place has been achieved .802</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1904838,
          "author_name": "Daisy",
          "author_url": "",
          "post_date": "2022-08-18T14:35:14.723000",
          "content": "<p>Yah, I also achieved .799 today, sorry for being late.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1904858,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-18T14:47:22.160000",
          "content": "<p>It's okay my friend. We live in different time zones and as long as we're still on the forum, we're never late.Good luck to you!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1892129,
      "author_name": "Robert Hatch",
      "author_url": "",
      "post_date": "2022-08-09T21:39:59.957000",
      "content": "<p>It's also possible we would've noticed gradual continual improvement if 4 digits were being reported. </p>\n<p>Without the fourth digit it's impossible to speculate how possible 802 is. Maybe the top teams would have a decent guess based on their own CV scores, but most of us have zero info to go on. The top score might be. 8014, but could also be .80199. </p>\n<p>We can guess 803 is looking super unlikely though ;). Diminishing returns means that would be a long long long ways past 802. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1891813,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2022-08-09T16:57:23.763000",
      "content": "<p>This seems like a bottleneck! Hope someone breaks it!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1904083,
      "author_name": "Jiwei Liu",
      "author_url": "",
      "post_date": "2022-08-18T00:09:47.843000",
      "content": "<p>Obviously not :P</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1900527,
      "author_name": "Mingjie Wang",
      "author_url": "",
      "post_date": "2022-08-16T05:04:20.617000",
      "content": "<p>This is dream!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1900385,
      "author_name": "duncan wood",
      "author_url": "",
      "post_date": "2022-08-16T01:45:43.377000",
      "content": "<p>I think it really depends on the underlying probability distribution. </p>\n<p>let's imagine it was a simple bimodal and either it was a 5% chance of default or 95% </p>\n<p>then the metric maxes out at 80-85, just because there are a bunch of positives in the data from people with a 5% chance of default, your model would rate them low as it should but it reduces the overall possible score</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1894077,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-11T08:55:06.407000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1891620": "Hi. Can someone technically explain why all model comes to 0.801 and (or very closely)?\nIs that common in other competition?",
    "1891837": "It is not common in all competitions, but it happens when the best public model is so close to the best model overall. I think as a group we could have gotten a slightly better score had noise not been introduced into data, and if we knew the exact meaning of all features. But chances are very good that for this particular metrics, we already have a score that is as good as it gets. If you think about it, a huge amount of time and computer resources have been spent (and likely wasted) on improving a score on the third decimal place, which in real world would never happen. Prizes and medals are big incentives.",
    "1893459": "In this competetion, the cv score and the LB score basically correspond. In my experiments, this correspondence becomes less reliable after my cv is higher than 0.8005 when LB reached 0.801. I think LB 0.801 might not the bottleneck, but our models are stucked in local optima. But time is running out and hopefully someone can break 0.801.",
    "1891831": "Though I think 0.801 is my limit, I feel like 0.802 is possible, considering that there are so many things that can be done...",
    "1912599": "Anonymized data necessarily reduces the accuracy of the predictions. Look at this notebook's explanation about B_19 and S_13. https://www.kaggle.com/code/ambrosm/amex-eda-which-makes-sense/notebook  (below cell 17) \n\nThe @roberthatch point about diminishing returns seems poignant. Was happy to see someone break .802",
    "1904506": "As of now, public lb first place has been achieved .802",
    "1892129": "It's also possible we would've noticed gradual continual improvement if 4 digits were being reported. \n\nWithout the fourth digit it's impossible to speculate how possible 802 is. Maybe the top teams would have a decent guess based on their own CV scores, but most of us have zero info to go on. The top score might be. 8014, but could also be .80199. \n\nWe can guess 803 is looking super unlikely though ;). Diminishing returns means that would be a long long long ways past 802. ",
    "1891813": "This seems like a bottleneck! Hope someone breaks it!",
    "1904083": "Obviously not :P",
    "1900527": "This is dream!",
    "1900385": "I think it really depends on the underlying probability distribution. \n\nlet's imagine it was a simple bimodal and either it was a 5% chance of default or 95% \n\nthen the metric maxes out at 80-85, just because there are a bunch of positives in the data from people with a 5% chance of default, your model would rate them low as it should but it reduces the overall possible score\n\n\n\n",
    "1894077": ""
  }
}