{
  "id": 347805,
  "title": "Takeaways from the competition…",
  "url": "/competitions/amex-default-prediction/discussion/347805",
  "author_name": "",
  "post_date": "2022-08-25T13:12:31.559311600Z",
  "votes": 7,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Beware, this is not a solutions writeup just ones log and takeaways. 🙂</p>\n<p>Takeaways from the competition…knowledge first and foremost! Alright I had some submissions in the silver zone but the choice went to others, wrong picking, but I know I’m not alone. Picking the best private sub sometimes happens but sometimes not, just how it is 😊</p>\n<p>Started working with the competition relative late, as the feature engineering is a big part, I started using some different public ds. Thanks to all the contribution, true teamwork within the race of score, the beauty of the Kaggle community.</p>\n<p>I thought I should take this competition to update the knowledge in new versions and the features in the automl area, so I did. <br>\nDidn’t manage to go trough the whole list but learned what was new in some. So the competition went more like a course than a competition, and I’m still running some in now the late submissions period, so the own created course didn’t stop at the deadline 😉 <br>\nWill not make a review of the findings, you know many of them already and for a fair public comparison it should be done another way, this was just for the own knowledge.<br>\nI didn’t used tuning and other advanced features in the different frameworks as that should take more time than I had. So getting an automl solution that had better score than some standalone tuned public versions and own standalone tuned, which I also used for the final selection, that wasn’t the purpose this time, but findings can be of value forward, creating a capital for some next challenge ahead or just for knowledge what’s out there right now. The AI train is going fast….</p>\n<p>So the key point here, sometimes it’s not only for the winning, it’s mostly for the knowledge, and for that I’m thankful! Thanks to the host, Kaggle and fellow Kagglers who made it possible. 🙏</p>\n<p>Also in this competition some ensemble techniques could be tested, when the different solutions where created. For me below went slightly better than basic mean ensemble.</p>\n<pre><code>import pandas as pd\nimport numpy as np\nimport scipy as sc\nimport os, random\nfrom collections import Counter\nfrom tqdm.auto import tqdm\n\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\nimport glob\n\nfrom scipy.stats import describe\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\nLABELS = [\"prediction\"]\n\nall_files2 = glob.glob(\"data/*.csv\")\nall_files2\n\nouts = [pd.read_csv(f, index_col=0) for f in all_files2]\n[f.sort_values(by='customer_ID', ascending=True, inplace=True)for f in outs]\nconcat_sub = pd.concat(outs, axis=1)\ncols = list(map(lambda x: \"m\" + str(x), range(len(concat_sub.columns))))\nconcat_sub.columns = cols\nconcat_sub.reset_index(inplace=True)\nconcat_sub = concat_sub.fillna('0')\n\ncorr = concat_sub.iloc[:,1:].corr()\nmask = np.zeros_like(corr, dtype=np.bool)\nmask[np.triu_indices_from(mask)] = True\n\n\nf, ax = plt.subplots(figsize=(len(cols)+2, len(cols)+2))\n\nsns.heatmap(corr,mask=mask,cmap='prism',vmin=0.95,center=0,linewidths=1,annot=True,fmt='.4f')\n\nrank = np.tril(concat_sub.iloc[:,1:].corr().values,-1)\nm = (rank&gt;0).sum()\nm_gmean, s = 0, 0\nfor n in range(min(rank.shape[0],m)):\n    mx = np.unravel_index(rank.argmin(), rank.shape)\n    w = (m-n)/(m+n)\n    print(w)\n    m_gmean += w*(np.log(concat_sub.iloc[:,mx[0]+1])+np.log(concat_sub.iloc[:,mx[1]+1]))/2\n    s += w\n    rank[mx] = 1\nm_gmean = np.exp(m_gmean/s)\n\nconcat_sub['prediction'] = m_gmean\nconcat_sub[['customer_ID','prediction']].to_csv('submission.csv', \n                                        index=False, float_format='%.4g')\n</code></pre>",
  "messages": [
    {
      "id": "1913714",
      "postDate": "08/25/2022 13:12:31",
      "content": "<p>Beware, this is not a solutions writeup just ones log and takeaways. 🙂</p>\n<p>Takeaways from the competition…knowledge first and foremost! Alright I had some submissions in the silver zone but the choice went to others, wrong picking, but I know I’m not alone. Picking the best private sub sometimes happens but sometimes not, just how it is 😊</p>\n<p>Started working with the competition relative late, as the feature engineering is a big part, I started using some different public ds. Thanks to all the contribution, true teamwork within the race of score, the beauty of the Kaggle community.</p>\n<p>I thought I should take this competition to update the knowledge in new versions and the features in the automl area, so I did. <br>\nDidn’t manage to go trough the whole list but learned what was new in some. So the competition went more like a course than a competition, and I’m still running some in now the late submissions period, so the own created course didn’t stop at the deadline 😉 <br>\nWill not make a review of the findings, you know many of them already and for a fair public comparison it should be done another way, this was just for the own knowledge.<br>\nI didn’t used tuning and other advanced features in the different frameworks as that should take more time than I had. So getting an automl solution that had better score than some standalone tuned public versions and own standalone tuned, which I also used for the final selection, that wasn’t the purpose this time, but findings can be of value forward, creating a capital for some next challenge ahead or just for knowledge what’s out there right now. The AI train is going fast….</p>\n<p>So the key point here, sometimes it’s not only for the winning, it’s mostly for the knowledge, and for that I’m thankful! Thanks to the host, Kaggle and fellow Kagglers who made it possible. 🙏</p>\n<p>Also in this competition some ensemble techniques could be tested, when the different solutions where created. For me below went slightly better than basic mean ensemble.</p>\n<pre><code>import pandas as pd\nimport numpy as np\nimport scipy as sc\nimport os, random\nfrom collections import Counter\nfrom tqdm.auto import tqdm\n\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\nimport glob\n\nfrom scipy.stats import describe\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\nLABELS = [\"prediction\"]\n\nall_files2 = glob.glob(\"data/*.csv\")\nall_files2\n\nouts = [pd.read_csv(f, index_col=0) for f in all_files2]\n[f.sort_values(by='customer_ID', ascending=True, inplace=True)for f in outs]\nconcat_sub = pd.concat(outs, axis=1)\ncols = list(map(lambda x: \"m\" + str(x), range(len(concat_sub.columns))))\nconcat_sub.columns = cols\nconcat_sub.reset_index(inplace=True)\nconcat_sub = concat_sub.fillna('0')\n\ncorr = concat_sub.iloc[:,1:].corr()\nmask = np.zeros_like(corr, dtype=np.bool)\nmask[np.triu_indices_from(mask)] = True\n\n\nf, ax = plt.subplots(figsize=(len(cols)+2, len(cols)+2))\n\nsns.heatmap(corr,mask=mask,cmap='prism',vmin=0.95,center=0,linewidths=1,annot=True,fmt='.4f')\n\nrank = np.tril(concat_sub.iloc[:,1:].corr().values,-1)\nm = (rank&gt;0).sum()\nm_gmean, s = 0, 0\nfor n in range(min(rank.shape[0],m)):\n    mx = np.unravel_index(rank.argmin(), rank.shape)\n    w = (m-n)/(m+n)\n    print(w)\n    m_gmean += w*(np.log(concat_sub.iloc[:,mx[0]+1])+np.log(concat_sub.iloc[:,mx[1]+1]))/2\n    s += w\n    rank[mx] = 1\nm_gmean = np.exp(m_gmean/s)\n\nconcat_sub['prediction'] = m_gmean\nconcat_sub[['customer_ID','prediction']].to_csv('submission.csv', \n                                        index=False, float_format='%.4g')\n</code></pre>",
      "rawMarkdown": "Beware, this is not a solutions writeup just ones log and takeaways. 🙂\n\nTakeaways from the competition…knowledge first and foremost! Alright I had some submissions in the silver zone but the choice went to others, wrong picking, but I know I’m not alone. Picking the best private sub sometimes happens but sometimes not, just how it is 😊\n\nStarted working with the competition relative late, as the feature engineering is a big part, I started using some different public ds. Thanks to all the contribution, true teamwork within the race of score, the beauty of the Kaggle community.\n\nI thought I should take this competition to update the knowledge in new versions and the features in the automl area, so I did. \nDidn’t manage to go trough the whole list but learned what was new in some. So the competition went more like a course than a competition, and I’m still running some in now the late submissions period, so the own created course didn’t stop at the deadline 😉 \nWill not make a review of the findings, you know many of them already and for a fair public comparison it should be done another way, this was just for the own knowledge.\nI didn’t used tuning and other advanced features in the different frameworks as that should take more time than I had. So getting an automl solution that had better score than some standalone tuned public versions and own standalone tuned, which I also used for the final selection, that wasn’t the purpose this time, but findings can be of value forward, creating a capital for some next challenge ahead or just for knowledge what’s out there right now. The AI train is going fast….\n\nSo the key point here, sometimes it’s not only for the winning, it’s mostly for the knowledge, and for that I’m thankful! Thanks to the host, Kaggle and fellow Kagglers who made it possible. 🙏\n\nAlso in this competition some ensemble techniques could be tested, when the different solutions where created. For me below went slightly better than basic mean ensemble.\n\n```\nimport pandas as pd\nimport numpy as np\nimport scipy as sc\nimport os, random\nfrom collections import Counter\nfrom tqdm.auto import tqdm\n\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\nimport glob\n\nfrom scipy.stats import describe\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\nLABELS = [\"prediction\"]\n\nall_files2 = glob.glob(\"data/*.csv\")\nall_files2\n\nouts = [pd.read_csv(f, index_col=0) for f in all_files2]\n[f.sort_values(by='customer_ID', ascending=True, inplace=True)for f in outs]\nconcat_sub = pd.concat(outs, axis=1)\ncols = list(map(lambda x: \"m\" + str(x), range(len(concat_sub.columns))))\nconcat_sub.columns = cols\nconcat_sub.reset_index(inplace=True)\nconcat_sub = concat_sub.fillna('0')\n\ncorr = concat_sub.iloc[:,1:].corr()\nmask = np.zeros_like(corr, dtype=np.bool)\nmask[np.triu_indices_from(mask)] = True\n\n\nf, ax = plt.subplots(figsize=(len(cols)+2, len(cols)+2))\n\nsns.heatmap(corr,mask=mask,cmap='prism',vmin=0.95,center=0,linewidths=1,annot=True,fmt='.4f')\n\nrank = np.tril(concat_sub.iloc[:,1:].corr().values,-1)\nm = (rank>0).sum()\nm_gmean, s = 0, 0\nfor n in range(min(rank.shape[0],m)):\n    mx = np.unravel_index(rank.argmin(), rank.shape)\n    w = (m-n)/(m+n)\n    print(w)\n    m_gmean += w*(np.log(concat_sub.iloc[:,mx[0]+1])+np.log(concat_sub.iloc[:,mx[1]+1]))/2\n    s += w\n    rank[mx] = 1\nm_gmean = np.exp(m_gmean/s)\n\nconcat_sub['prediction'] = m_gmean\nconcat_sub[['customer_ID','prediction']].to_csv('submission.csv', \n                                        index=False, float_format='%.4g')\n```",
      "votes": null
    },
    {
      "id": "1914639",
      "postDate": "08/26/2022 08:56:01",
      "content": "<p>What is your best score for autoML and which autoML tool ?<br>\nMy autogluon model gave me public 0.797  (private : 0.80524)</p>",
      "rawMarkdown": "What is your best score for autoML and which autoML tool ?\nMy autogluon model gave me public 0.797  (private : 0.80524)",
      "votes": null
    },
    {
      "id": "1914644",
      "postDate": "08/26/2022 09:07:29",
      "content": "<p>Haven't tried them all yet, still running :) With autogluon and with only standard setup I got 0.80507 in private. I guess one can squeeze more out of it with hyp tuning and more features, they have destill and pseudo options for example.</p>",
      "rawMarkdown": "Haven't tried them all yet, still running :) With autogluon and with only standard setup I got 0.80507 in private. I guess one can squeeze more out of it with hyp tuning and more features, they have destill and pseudo options for example.",
      "votes": null
    },
    {
      "id": "1914648",
      "postDate": "08/26/2022 09:12:07",
      "content": "<p>Almost all of the standard AutoMLs have rapids now included as feature, good updates!</p>",
      "rawMarkdown": "Almost all of the standard AutoMLs have rapids now included as feature, good updates!",
      "votes": null
    },
    {
      "id": "1922122",
      "postDate": "09/01/2022 09:28:34",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kirderf\" target=\"_blank\">@kirderf</a> May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "rawMarkdown": "Hi @kirderf May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1914639,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "08/26/2022 08:56:01",
      "content": "<p>What is your best score for autoML and which autoML tool ?<br>\nMy autogluon model gave me public 0.797  (private : 0.80524)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1914644,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "08/26/2022 09:07:29",
          "content": "<p>Haven't tried them all yet, still running :) With autogluon and with only standard setup I got 0.80507 in private. I guess one can squeeze more out of it with hyp tuning and more features, they have destill and pseudo options for example.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1914648,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "08/26/2022 09:12:07",
          "content": "<p>Almost all of the standard AutoMLs have rapids now included as feature, good updates!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1922122,
      "author_name": "lystriving",
      "author_url": "",
      "post_date": "09/01/2022 09:28:34",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kirderf\" target=\"_blank\">@kirderf</a> May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: <a href=\"https://cityuhk.questionpro.com/survey-of-kaggle-contestants\" target=\"_blank\">https://cityuhk.questionpro.com/survey-of-kaggle-contestants</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1913714": "Beware, this is not a solutions writeup just ones log and takeaways. 🙂\n\nTakeaways from the competition…knowledge first and foremost! Alright I had some submissions in the silver zone but the choice went to others, wrong picking, but I know I’m not alone. Picking the best private sub sometimes happens but sometimes not, just how it is 😊\n\nStarted working with the competition relative late, as the feature engineering is a big part, I started using some different public ds. Thanks to all the contribution, true teamwork within the race of score, the beauty of the Kaggle community.\n\nI thought I should take this competition to update the knowledge in new versions and the features in the automl area, so I did. \nDidn’t manage to go trough the whole list but learned what was new in some. So the competition went more like a course than a competition, and I’m still running some in now the late submissions period, so the own created course didn’t stop at the deadline 😉 \nWill not make a review of the findings, you know many of them already and for a fair public comparison it should be done another way, this was just for the own knowledge.\nI didn’t used tuning and other advanced features in the different frameworks as that should take more time than I had. So getting an automl solution that had better score than some standalone tuned public versions and own standalone tuned, which I also used for the final selection, that wasn’t the purpose this time, but findings can be of value forward, creating a capital for some next challenge ahead or just for knowledge what’s out there right now. The AI train is going fast….\n\nSo the key point here, sometimes it’s not only for the winning, it’s mostly for the knowledge, and for that I’m thankful! Thanks to the host, Kaggle and fellow Kagglers who made it possible. 🙏\n\nAlso in this competition some ensemble techniques could be tested, when the different solutions where created. For me below went slightly better than basic mean ensemble.\n\n```\nimport pandas as pd\nimport numpy as np\nimport scipy as sc\nimport os, random\nfrom collections import Counter\nfrom tqdm.auto import tqdm\n\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\nimport glob\n\nfrom scipy.stats import describe\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\nLABELS = [\"prediction\"]\n\nall_files2 = glob.glob(\"data/*.csv\")\nall_files2\n\nouts = [pd.read_csv(f, index_col=0) for f in all_files2]\n[f.sort_values(by='customer_ID', ascending=True, inplace=True)for f in outs]\nconcat_sub = pd.concat(outs, axis=1)\ncols = list(map(lambda x: \"m\" + str(x), range(len(concat_sub.columns))))\nconcat_sub.columns = cols\nconcat_sub.reset_index(inplace=True)\nconcat_sub = concat_sub.fillna('0')\n\ncorr = concat_sub.iloc[:,1:].corr()\nmask = np.zeros_like(corr, dtype=np.bool)\nmask[np.triu_indices_from(mask)] = True\n\n\nf, ax = plt.subplots(figsize=(len(cols)+2, len(cols)+2))\n\nsns.heatmap(corr,mask=mask,cmap='prism',vmin=0.95,center=0,linewidths=1,annot=True,fmt='.4f')\n\nrank = np.tril(concat_sub.iloc[:,1:].corr().values,-1)\nm = (rank>0).sum()\nm_gmean, s = 0, 0\nfor n in range(min(rank.shape[0],m)):\n    mx = np.unravel_index(rank.argmin(), rank.shape)\n    w = (m-n)/(m+n)\n    print(w)\n    m_gmean += w*(np.log(concat_sub.iloc[:,mx[0]+1])+np.log(concat_sub.iloc[:,mx[1]+1]))/2\n    s += w\n    rank[mx] = 1\nm_gmean = np.exp(m_gmean/s)\n\nconcat_sub['prediction'] = m_gmean\nconcat_sub[['customer_ID','prediction']].to_csv('submission.csv', \n                                        index=False, float_format='%.4g')\n```",
    "1914639": "What is your best score for autoML and which autoML tool ?\nMy autogluon model gave me public 0.797  (private : 0.80524)",
    "1914644": "Haven't tried them all yet, still running :) With autogluon and with only standard setup I got 0.80507 in private. I guess one can squeeze more out of it with hyp tuning and more features, they have destill and pseudo options for example.",
    "1914648": "Almost all of the standard AutoMLs have rapids now included as feature, good updates!",
    "1922122": "Hi @kirderf May I invite you to participate in this survey regarding your experience on Kaggle (10 min)? This is not a scam. We are a group of researchers at the City University of Hong Kong. The survey link is: https://cityuhk.questionpro.com/survey-of-kaggle-contestants"
  },
  "source": "meta"
}