{
  "id": 336146,
  "title": "Ranking approaches ?",
  "url": "/competitions/amex-default-prediction/discussion/336146",
  "author_name": "Lucas Morin",
  "post_date": "2022-07-09T15:35:46.574000",
  "votes": 13,
  "comment_count": 8,
  "views": 0,
  "content": "<p>As the metrics are inherently relative ranking metrics I tought it might be a good idea to approach this problem as a ranking problem. However, gbdt seems to have a cap on query size. And I wasn't able to find a good nn starter. Has there already been a tabular ranking competition ?  Anyone has tried this approach here / is willing to share a baseline ?</p>",
  "messages": [
    {
      "id": 1849504,
      "postDate": "2022-07-09T15:35:46.573Z",
      "content": "<p>As the metrics are inherently relative ranking metrics I tought it might be a good idea to approach this problem as a ranking problem. However, gbdt seems to have a cap on query size. And I wasn't able to find a good nn starter. Has there already been a tabular ranking competition ?  Anyone has tried this approach here / is willing to share a baseline ?</p>",
      "rawMarkdown": "As the metrics are inherently relative ranking metrics I tought it might be a good idea to approach this problem as a ranking problem. However, gbdt seems to have a cap on query size. And I wasn't able to find a good nn starter. Has there already been a tabular ranking competition ?  Anyone has tried this approach here / is willing to share a baseline ?",
      "votes": 13
    },
    {
      "id": 1851759,
      "postDate": "2022-07-11T14:01:05.017Z",
      "content": "<p>I am extremely happy with 'objective' : 'rank:pairwise', and eval of auc, in XGB model. </p>\n<p>It should work fine drop in to any XGB model, but for more concrete example, I'm also hoping to publicly share my CV 0.7963 XGB notebook if I can finish it up soon. </p>",
      "rawMarkdown": "I am extremely happy with 'objective' : 'rank:pairwise', and eval of auc, in XGB model. \n\nIt should work fine drop in to any XGB model, but for more concrete example, I'm also hoping to publicly share my CV 0.7963 XGB notebook if I can finish it up soon. ",
      "votes": 3,
      "replies": [
        {
          "id": 1852699,
          "postDate": "2022-07-12T09:19:40.033Z",
          "content": "<p>Ha yes it seems that XGB 'rank:pairwise' doesn't have the same constraints as lgbm. Very nice of you to share. I am still hoping to find baseline NN architecture for ranking on tabular data. </p>",
          "rawMarkdown": "Ha yes it seems that XGB 'rank:pairwise' doesn't have the same constraints as lgbm. Very nice of you to share. I am still hoping to find baseline NN architecture for ranking on tabular data. "
        },
        {
          "id": 1853924,
          "postDate": "2022-07-13T09:52:58.680Z",
          "content": "<p>Hi Robert, what do you do with the output of \"rank:pairwise\" on XGB to probability?</p>",
          "rawMarkdown": "Hi Robert, what do you do with the output of \"rank:pairwise\" on XGB to probability?"
        },
        {
          "id": 1853935,
          "postDate": "2022-07-13T10:10:34.483Z",
          "content": "<p>We have ranking metrics, no need to have well calibrated models for submission. (calibration in probility might be needed for ensembling …)</p>",
          "rawMarkdown": "We have ranking metrics, no need to have well calibrated models for submission. (calibration in probility might be needed for ensembling ...)",
          "votes": 1
        },
        {
          "id": 1854297,
          "postDate": "2022-07-13T15:32:40.693Z",
          "content": "<p>Use raw output for single model or blend of exact same models. For ensemble, I would convert (all models, not just this one) to ranking and add them together. Ranking ensembles is one of many topics covered in depth in that ensemble primer linked in this forum here: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/332729\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/332729</a></p>",
          "rawMarkdown": "Use raw output for single model or blend of exact same models. For ensemble, I would convert (all models, not just this one) to ranking and add them together. Ranking ensembles is one of many topics covered in depth in that ensemble primer linked in this forum here: https://www.kaggle.com/competitions/amex-default-prediction/discussion/332729",
          "votes": 2
        }
      ]
    },
    {
      "id": 1853640,
      "postDate": "2022-07-13T03:10:03.270Z",
      "content": "<p>I tried bpe loss along with the cross entropy loss in my gru model. <a href=\"https://www.kaggle.com/code/narendra/amex-gru-rank-train/notebook\" target=\"_blank\">here</a>.<br>\nIt gave similar cv and lb score compared with the model without bpe loss. But i see it as a regularizer(since adding constraint), cv is going up smoothly and was able to run longer iterations.</p>",
      "rawMarkdown": "I tried bpe loss along with the cross entropy loss in my gru model. [here](https://www.kaggle.com/code/narendra/amex-gru-rank-train/notebook).\nIt gave similar cv and lb score compared with the model without bpe loss. But i see it as a regularizer(since adding constraint), cv is going up smoothly and was able to run longer iterations."
    },
    {
      "id": 1850359,
      "postDate": "2022-07-10T10:51:44.237Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1850369,
          "postDate": "2022-07-10T10:59:28.803Z",
          "content": "<p>Yeah there is all sorts of things to try on the feature engineering side. Esp. for the features with shifting distributions. I was wondering about the modelling aspects of outputing a rank or something assimilated to it.</p>",
          "rawMarkdown": "Yeah there is all sorts of things to try on the feature engineering side. Esp. for the features with shifting distributions. I was wondering about the modelling aspects of outputing a rank or something assimilated to it."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1851759,
      "author_name": "Robert Hatch",
      "author_url": "",
      "post_date": "2022-07-11T14:01:05.017000",
      "content": "<p>I am extremely happy with 'objective' : 'rank:pairwise', and eval of auc, in XGB model. </p>\n<p>It should work fine drop in to any XGB model, but for more concrete example, I'm also hoping to publicly share my CV 0.7963 XGB notebook if I can finish it up soon. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 1852699,
          "author_name": "Lucas Morin",
          "author_url": "",
          "post_date": "2022-07-12T09:19:40.033000",
          "content": "<p>Ha yes it seems that XGB 'rank:pairwise' doesn't have the same constraints as lgbm. Very nice of you to share. I am still hoping to find baseline NN architecture for ranking on tabular data. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1853924,
          "author_name": "Jose Antonio Alatorre",
          "author_url": "",
          "post_date": "2022-07-13T09:52:58.680000",
          "content": "<p>Hi Robert, what do you do with the output of \"rank:pairwise\" on XGB to probability?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1853935,
          "author_name": "Lucas Morin",
          "author_url": "",
          "post_date": "2022-07-13T10:10:34.483000",
          "content": "<p>We have ranking metrics, no need to have well calibrated models for submission. (calibration in probility might be needed for ensembling …)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1854297,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2022-07-13T15:32:40.693000",
          "content": "<p>Use raw output for single model or blend of exact same models. For ensemble, I would convert (all models, not just this one) to ranking and add them together. Ranking ensembles is one of many topics covered in depth in that ensemble primer linked in this forum here: <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/332729\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/332729</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1853640,
      "author_name": "doteeee",
      "author_url": "",
      "post_date": "2022-07-13T03:10:03.270000",
      "content": "<p>I tried bpe loss along with the cross entropy loss in my gru model. <a href=\"https://www.kaggle.com/code/narendra/amex-gru-rank-train/notebook\" target=\"_blank\">here</a>.<br>\nIt gave similar cv and lb score compared with the model without bpe loss. But i see it as a regularizer(since adding constraint), cv is going up smoothly and was able to run longer iterations.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1850359,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-10T10:51:44.237000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1850369,
          "author_name": "Lucas Morin",
          "author_url": "",
          "post_date": "2022-07-10T10:59:28.803000",
          "content": "<p>Yeah there is all sorts of things to try on the feature engineering side. Esp. for the features with shifting distributions. I was wondering about the modelling aspects of outputing a rank or something assimilated to it.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1849504": "As the metrics are inherently relative ranking metrics I tought it might be a good idea to approach this problem as a ranking problem. However, gbdt seems to have a cap on query size. And I wasn't able to find a good nn starter. Has there already been a tabular ranking competition ?  Anyone has tried this approach here / is willing to share a baseline ?",
    "1851759": "I am extremely happy with 'objective' : 'rank:pairwise', and eval of auc, in XGB model. \n\nIt should work fine drop in to any XGB model, but for more concrete example, I'm also hoping to publicly share my CV 0.7963 XGB notebook if I can finish it up soon. ",
    "1853640": "I tried bpe loss along with the cross entropy loss in my gru model. [here](https://www.kaggle.com/code/narendra/amex-gru-rank-train/notebook).\nIt gave similar cv and lb score compared with the model without bpe loss. But i see it as a regularizer(since adding constraint), cv is going up smoothly and was able to run longer iterations.",
    "1850359": ""
  }
}