{
  "id": 381630,
  "title": "Can we make the miracle with larger computation resources?",
  "url": "/competitions/otto-recommender-system/discussion/381630",
  "author_name": "",
  "post_date": "2023-01-27T13:12:33.825584100Z",
  "votes": 6,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I found most people use rule-based methods to select topN candidates for each session to rank, e.g. 200 candidates. Then the final score would also be limited by the topN selection stage. How about selecting the top 800 or even &gt; 1000 candidates for ranking? Can it get a better result?</p>",
  "messages": [
    {
      "id": "2117693",
      "postDate": "01/27/2023 13:12:33",
      "content": "<p>I found most people use rule-based methods to select topN candidates for each session to rank, e.g. 200 candidates. Then the final score would also be limited by the topN selection stage. How about selecting the top 800 or even &gt; 1000 candidates for ranking? Can it get a better result?</p>",
      "rawMarkdown": "I found most people use rule-based methods to select topN candidates for each session to rank, e.g. 200 candidates. Then the final score would also be limited by the topN selection stage. How about selecting the top 800 or even > 1000 candidates for ranking? Can it get a better result?",
      "votes": null
    },
    {
      "id": "2117699",
      "postDate": "01/27/2023 13:14:47",
      "content": "<p>Increasing candidates doesn't improve my LB. I found that when I changed candidates number from 50 to 100, the LB decreased. Is this phenomenon normal? </p>",
      "rawMarkdown": "Increasing candidates doesn't improve my LB. I found that when I changed candidates number from 50 to 100, the LB decreased. Is this phenomenon normal?",
      "votes": null
    },
    {
      "id": "2117704",
      "postDate": "01/27/2023 13:18:33",
      "content": "<p>That depends on you rerank model's performance</p>",
      "rawMarkdown": "That depends on you rerank model's performance",
      "votes": null
    },
    {
      "id": "2117722",
      "postDate": "01/27/2023 13:48:52",
      "content": "<p>I tried to increase candidate count by a lot at once but it didn't work with my previous feature set. My theoretical max recall increased to 0.648 but model performance degraded a lot. This is my first recommender system project so I'm not quite familiar with those outcomes. I don't how to evaluate that but that's what happened to me.</p>",
      "rawMarkdown": "I tried to increase candidate count by a lot at once but it didn't work with my previous feature set. My theoretical max recall increased to 0.648 but model performance degraded a lot. This is my first recommender system project so I'm not quite familiar with those outcomes. I don't how to evaluate that but that's what happened to me.",
      "votes": null
    },
    {
      "id": "2117728",
      "postDate": "01/27/2023 13:54:06",
      "content": "<p>Does anyone notice improvement between using 50 100 or 200 candidates ??</p>",
      "rawMarkdown": "Does anyone notice improvement between using 50 100 or 200 candidates ??",
      "votes": null
    },
    {
      "id": "2117730",
      "postDate": "01/27/2023 13:56:53",
      "content": "<p>For my solution, it probably does not work.<br>\nBecause I tried to increase my retrival quality in the middle of the comp.<br>\nAnd it did work, the retrival rate increased by 0.005 under the same retrieved candidate number, which is considered a huge boost in my opition. (imagine from 0.603 to 0.608…).<br>\nHowever it turns out the final improvement is only 0.0004 based on my local CV.<br>\nIt means the retriver did do a better job and retrieved some very \"difficult\" samples, but they are too difficult for the ranker to pick them out.<br>\nSo, in your proposal, it would probably the similar case, the increased retrival number can retrieve more correct candidates for sure, but they are \"difficult\" samples and your ranker may be not strong enough to pick them out.<br>\nAnyway, I think it depends on your ranker capability.</p>",
      "rawMarkdown": "For my solution, it probably does not work.\nBecause I tried to increase my retrival quality in the middle of the comp.\nAnd it did work, the retrival rate increased by 0.005 under the same retrieved candidate number, which is considered a huge boost in my opition. (imagine from 0.603 to 0.608...).\nHowever it turns out the final improvement is only 0.0004 based on my local CV.\nIt means the retriver did do a better job and retrieved some very \"difficult\" samples, but they are too difficult for the ranker to pick them out.\nSo, in your proposal, it would probably the similar case, the increased retrival number can retrieve more correct candidates for sure, but they are \"difficult\" samples and your ranker may be not strong enough to pick them out.\nAnyway, I think it depends on your ranker capability.",
      "votes": null
    },
    {
      "id": "2117733",
      "postDate": "01/27/2023 13:57:55",
      "content": "<p>I used 200 candidates from the beginning. Currently, I'm trying to use 600 candidates. Trying to make the miracle with larger resources😜</p>",
      "rawMarkdown": "I used 200 candidates from the beginning. Currently, I'm trying to use 600 candidates. Trying to make the miracle with larger resources😜",
      "votes": null
    },
    {
      "id": "2117735",
      "postDate": "01/27/2023 14:00:17",
      "content": "<p>Thanks for sharing. Make sense. Just to give it a try. If it doesn't work by the end of tmr. We'll focus on ensembling</p>",
      "rawMarkdown": "Thanks for sharing. Make sense. Just to give it a try. If it doesn't work by the end of tmr. We'll focus on ensembling",
      "votes": null
    },
    {
      "id": "2117740",
      "postDate": "01/27/2023 14:10:52",
      "content": "<p>Well, do let me know the experiment result after the competition. Just curious and validate my thought.</p>",
      "rawMarkdown": "Well, do let me know the experiment result after the competition. Just curious and validate my thought.",
      "votes": null
    },
    {
      "id": "2117756",
      "postDate": "01/27/2023 14:16:01",
      "content": "<p>sure thing:)</p>",
      "rawMarkdown": "sure thing:)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2117699,
      "author_name": "kimoyami",
      "author_url": "",
      "post_date": "01/27/2023 13:14:47",
      "content": "<p>Increasing candidates doesn't improve my LB. I found that when I changed candidates number from 50 to 100, the LB decreased. Is this phenomenon normal? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2117704,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "01/27/2023 13:18:33",
          "content": "<p>That depends on you rerank model's performance</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2117722,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "01/27/2023 13:48:52",
      "content": "<p>I tried to increase candidate count by a lot at once but it didn't work with my previous feature set. My theoretical max recall increased to 0.648 but model performance degraded a lot. This is my first recommender system project so I'm not quite familiar with those outcomes. I don't how to evaluate that but that's what happened to me.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2117728,
      "author_name": "rayanaay",
      "author_url": "",
      "post_date": "01/27/2023 13:54:06",
      "content": "<p>Does anyone notice improvement between using 50 100 or 200 candidates ??</p>",
      "votes": null,
      "replies": [
        {
          "id": 2117733,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "01/27/2023 13:57:55",
          "content": "<p>I used 200 candidates from the beginning. Currently, I'm trying to use 600 candidates. Trying to make the miracle with larger resources😜</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2117730,
      "author_name": "buumoo",
      "author_url": "",
      "post_date": "01/27/2023 13:56:53",
      "content": "<p>For my solution, it probably does not work.<br>\nBecause I tried to increase my retrival quality in the middle of the comp.<br>\nAnd it did work, the retrival rate increased by 0.005 under the same retrieved candidate number, which is considered a huge boost in my opition. (imagine from 0.603 to 0.608…).<br>\nHowever it turns out the final improvement is only 0.0004 based on my local CV.<br>\nIt means the retriver did do a better job and retrieved some very \"difficult\" samples, but they are too difficult for the ranker to pick them out.<br>\nSo, in your proposal, it would probably the similar case, the increased retrival number can retrieve more correct candidates for sure, but they are \"difficult\" samples and your ranker may be not strong enough to pick them out.<br>\nAnyway, I think it depends on your ranker capability.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2117735,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "01/27/2023 14:00:17",
          "content": "<p>Thanks for sharing. Make sense. Just to give it a try. If it doesn't work by the end of tmr. We'll focus on ensembling</p>",
          "votes": null,
          "replies": [
            {
              "id": 2117740,
              "author_name": "buumoo",
              "author_url": "",
              "post_date": "01/27/2023 14:10:52",
              "content": "<p>Well, do let me know the experiment result after the competition. Just curious and validate my thought.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2117756,
                  "author_name": "wuwenmin",
                  "author_url": "",
                  "post_date": "01/27/2023 14:16:01",
                  "content": "<p>sure thing:)</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2117693": "I found most people use rule-based methods to select topN candidates for each session to rank, e.g. 200 candidates. Then the final score would also be limited by the topN selection stage. How about selecting the top 800 or even > 1000 candidates for ranking? Can it get a better result?",
    "2117699": "Increasing candidates doesn't improve my LB. I found that when I changed candidates number from 50 to 100, the LB decreased. Is this phenomenon normal?",
    "2117704": "That depends on you rerank model's performance",
    "2117722": "I tried to increase candidate count by a lot at once but it didn't work with my previous feature set. My theoretical max recall increased to 0.648 but model performance degraded a lot. This is my first recommender system project so I'm not quite familiar with those outcomes. I don't how to evaluate that but that's what happened to me.",
    "2117728": "Does anyone notice improvement between using 50 100 or 200 candidates ??",
    "2117730": "For my solution, it probably does not work.\nBecause I tried to increase my retrival quality in the middle of the comp.\nAnd it did work, the retrival rate increased by 0.005 under the same retrieved candidate number, which is considered a huge boost in my opition. (imagine from 0.603 to 0.608...).\nHowever it turns out the final improvement is only 0.0004 based on my local CV.\nIt means the retriver did do a better job and retrieved some very \"difficult\" samples, but they are too difficult for the ranker to pick them out.\nSo, in your proposal, it would probably the similar case, the increased retrival number can retrieve more correct candidates for sure, but they are \"difficult\" samples and your ranker may be not strong enough to pick them out.\nAnyway, I think it depends on your ranker capability.",
    "2117733": "I used 200 candidates from the beginning. Currently, I'm trying to use 600 candidates. Trying to make the miracle with larger resources😜",
    "2117735": "Thanks for sharing. Make sense. Just to give it a try. If it doesn't work by the end of tmr. We'll focus on ensembling",
    "2117740": "Well, do let me know the experiment result after the competition. Just curious and validate my thought.",
    "2117756": "sure thing:)"
  },
  "source": "meta"
}