{
  "id": 368564,
  "title": "Evaluation of stage1.Candidates",
  "url": "/competitions/otto-recommender-system/discussion/368564",
  "author_name": "",
  "post_date": "2022-11-26T08:12:40.496486600Z",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>As shown in the <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\" target=\"_blank\">discussion</a> here and in the <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">public code</a>, <br>\nIn the OTTO competition, 2-stage recommendation is baseline.</p>\n<p>stage1. candidate generation<br>\nstage2. ranking</p>\n<p>We would like to evaluate each candidate at stage1.<br>\nI am thinking of evaluating each candidate using recall@k, which is the metrix for this competition(we will probably use recall@20), but I am not sure if this is the appropriate way to do it.<br>\nI would be very grateful if you could give me your opinions.</p>\n<p>Candidate evaluation (draft)</p>\n<ul>\n<li>recall@k * Competition evaluation metrics. If candidate over 20, candidates must be ranked.</li>\n<li>precision@k * If candidate over 20, candidates must be ranked.</li>\n<li>Percentage of hits: number of hits/number of preds.<br>\netc.</li>\n</ul>\n<p>I think the following metrics is not appropriate for evaluation, because the higher the number of candidates, the higher the evaluation (if you predict all article candidate, recall/precision will be 1.0).</p>\n<ul>\n<li>recall</li>\n<li>precision</li>\n<li>hit counts<br>\netc.</li>\n</ul>",
  "messages": [
    {
      "id": "2044086",
      "postDate": "11/26/2022 08:12:40",
      "content": "<p>As shown in the <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\" target=\"_blank\">discussion</a> here and in the <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">public code</a>, <br>\nIn the OTTO competition, 2-stage recommendation is baseline.</p>\n<p>stage1. candidate generation<br>\nstage2. ranking</p>\n<p>We would like to evaluate each candidate at stage1.<br>\nI am thinking of evaluating each candidate using recall@k, which is the metrix for this competition(we will probably use recall@20), but I am not sure if this is the appropriate way to do it.<br>\nI would be very grateful if you could give me your opinions.</p>\n<p>Candidate evaluation (draft)</p>\n<ul>\n<li>recall@k * Competition evaluation metrics. If candidate over 20, candidates must be ranked.</li>\n<li>precision@k * If candidate over 20, candidates must be ranked.</li>\n<li>Percentage of hits: number of hits/number of preds.<br>\netc.</li>\n</ul>\n<p>I think the following metrics is not appropriate for evaluation, because the higher the number of candidates, the higher the evaluation (if you predict all article candidate, recall/precision will be 1.0).</p>\n<ul>\n<li>recall</li>\n<li>precision</li>\n<li>hit counts<br>\netc.</li>\n</ul>",
      "rawMarkdown": "As shown in the [discussion](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721) here and in the [public code](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575), \nIn the OTTO competition, 2-stage recommendation is baseline.\n\nstage1. candidate generation\nstage2. ranking\n\nWe would like to evaluate each candidate at stage1.\nI am thinking of evaluating each candidate using recall@k, which is the metrix for this competition(we will probably use recall@20), but I am not sure if this is the appropriate way to do it.\nI would be very grateful if you could give me your opinions.\n\nCandidate evaluation (draft)\n- recall@k * Competition evaluation metrics. If candidate over 20, candidates must be ranked.\n- precision@k * If candidate over 20, candidates must be ranked.\n- Percentage of hits: number of hits/number of preds.\netc.\n\nI think the following metrics is not appropriate for evaluation, because the higher the number of candidates, the higher the evaluation (if you predict all article candidate, recall/precision will be 1.0).\n- recall\n- precision\n- hit counts\netc.",
      "votes": null
    },
    {
      "id": "2115910",
      "postDate": "01/26/2023 04:16:10",
      "content": "<p>Hi. I think for candidate generation. Recall might be a better metric to evaluate.  The goal of candiadate generation is to extract all the candiates that we need, so it's better that the cadidates includes all the the groud truth. In this case, we can focus on ranker to rank it better</p>",
      "rawMarkdown": "Hi. I think for candidate generation. Recall might be a better metric to evaluate.  The goal of candiadate generation is to extract all the candiates that we need, so it's better that the cadidates includes all the the groud truth. In this case, we can focus on ranker to rank it better",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2115910,
      "author_name": "huaguo",
      "author_url": "",
      "post_date": "01/26/2023 04:16:10",
      "content": "<p>Hi. I think for candidate generation. Recall might be a better metric to evaluate.  The goal of candiadate generation is to extract all the candiates that we need, so it's better that the cadidates includes all the the groud truth. In this case, we can focus on ranker to rank it better</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2044086": "As shown in the [discussion](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721) here and in the [public code](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575), \nIn the OTTO competition, 2-stage recommendation is baseline.\n\nstage1. candidate generation\nstage2. ranking\n\nWe would like to evaluate each candidate at stage1.\nI am thinking of evaluating each candidate using recall@k, which is the metrix for this competition(we will probably use recall@20), but I am not sure if this is the appropriate way to do it.\nI would be very grateful if you could give me your opinions.\n\nCandidate evaluation (draft)\n- recall@k * Competition evaluation metrics. If candidate over 20, candidates must be ranked.\n- precision@k * If candidate over 20, candidates must be ranked.\n- Percentage of hits: number of hits/number of preds.\netc.\n\nI think the following metrics is not appropriate for evaluation, because the higher the number of candidates, the higher the evaluation (if you predict all article candidate, recall/precision will be 1.0).\n- recall\n- precision\n- hit counts\netc.",
    "2115910": "Hi. I think for candidate generation. Recall might be a better metric to evaluate.  The goal of candiadate generation is to extract all the candiates that we need, so it's better that the cadidates includes all the the groud truth. In this case, we can focus on ranker to rank it better"
  },
  "source": "meta"
}