{
  "id": 412346,
  "title": "How to adjust the best threshold",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/412346",
  "author_name": "",
  "post_date": "2023-05-23T10:16:02.348755700Z",
  "votes": 2,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Many notebooks set the same threshold for all questions, but I don't think it is the best.<br>\nWhen I set the thresholds for Q1 to Q18 individually, the f1-score increased by about 0.003(LB).<br>\nBut CV doesn't do the same, so it may be overfitting LB.<br>\nWhat would you do?</p>",
  "messages": [
    {
      "id": "2270692",
      "postDate": "05/23/2023 10:16:02",
      "content": "<p>Many notebooks set the same threshold for all questions, but I don't think it is the best.<br>\nWhen I set the thresholds for Q1 to Q18 individually, the f1-score increased by about 0.003(LB).<br>\nBut CV doesn't do the same, so it may be overfitting LB.<br>\nWhat would you do?</p>",
      "rawMarkdown": "Many notebooks set the same threshold for all questions, but I don't think it is the best.\nWhen I set the thresholds for Q1 to Q18 individually, the f1-score increased by about 0.003(LB).\nBut CV doesn't do the same, so it may be overfitting LB.\nWhat would you do?",
      "votes": null
    },
    {
      "id": "2270698",
      "postDate": "05/23/2023 10:23:12",
      "content": "<p>good idea. i also think this idea. but i only try 3 group 'for' code ,it largely waste time. how do you try q1 to q18 ? use for and other?</p>",
      "rawMarkdown": "good idea. i also think this idea. but i only try 3 group 'for' code ,it largely waste time. how do you try q1 to q18 ? use for and other?",
      "votes": null
    },
    {
      "id": "2270701",
      "postDate": "05/23/2023 10:26:04",
      "content": "<p>I think so, too.</p>\n<p>How did you decide on the threshold from Q1 to Q18 individually?</p>",
      "rawMarkdown": "I think so, too.\n\nHow did you decide on the threshold from Q1 to Q18 individually?",
      "votes": null
    },
    {
      "id": "2270824",
      "postDate": "05/23/2023 13:02:51",
      "content": "<p>I created a list containing 18 thresholds, and executed it with 'for' code.<br>\nI changed these threshold little by little and repeated submit.<br>\nIt's not very smart method.</p>",
      "rawMarkdown": "I created a list containing 18 thresholds, and executed it with 'for' code.\nI changed these threshold little by little and repeated submit.\nIt's not very smart method.",
      "votes": null
    },
    {
      "id": "2270833",
      "postDate": "05/23/2023 13:05:48",
      "content": "<p>I changed thresholds little by little and repeated submissions.<br>\nI don't know of an efficient way, sorry.</p>",
      "rawMarkdown": "I changed thresholds little by little and repeated submissions.\nI don't know of an efficient way, sorry.",
      "votes": null
    },
    {
      "id": "2270844",
      "postDate": "05/23/2023 13:10:26",
      "content": "<p>thanks, i want to share my idea with you. </p>\n<pre><code>scores = []; thresholds = []\nbest_score = 0; best_threshold1 = 0\n\nfor i in np.arange(0.60,0.65,0.005):\n    for j in np.arange(0.60,0.65,0.005):\n        for k in np.arange(0.60,0.65,0.005):\n            # print(f'{i}{j}{k}',end='')\n            preds=pd.concat([(oof[[0,1,2]]&gt;i).astype('int'),(oof[[3,4,5,6,7,8,9,10,11,12]]&gt;j).astype('int'),(oof[[13,14,15,16,17]]&gt;k).astype('int')],axis=1)\n\n            m = f1_score(true.to_pandas().values.reshape((-1)), preds.to_pandas().values.reshape((-1)), average='macro')   \n            scores.append(m)\n            thresholds.append(threshold)\n            if m&gt;best_score:\n                best_score = m\n\n                best_threshold1 = (i,j,k)\n                # print(best_score,best_threshold)\nprint()\ndisplay(best_threshold1,best_score)\n</code></pre>\n<p>it is a slowly idea . then i will try parallel to accelerate this code or try 18 'for' code (It's also not very smart method.)</p>",
      "rawMarkdown": "thanks, i want to share my idea with you. \n\n\n```\nscores = []; thresholds = []\nbest_score = 0; best_threshold1 = 0\n\nfor i in np.arange(0.60,0.65,0.005):\n    for j in np.arange(0.60,0.65,0.005):\n        for k in np.arange(0.60,0.65,0.005):\n            # print(f'{i}{j}{k}',end='')\n            preds=pd.concat([(oof[[0,1,2]]>i).astype('int'),(oof[[3,4,5,6,7,8,9,10,11,12]]>j).astype('int'),(oof[[13,14,15,16,17]]>k).astype('int')],axis=1)\n\n            m = f1_score(true.to_pandas().values.reshape((-1)), preds.to_pandas().values.reshape((-1)), average='macro')   \n            scores.append(m)\n            thresholds.append(threshold)\n            if m>best_score:\n                best_score = m\n                \n                best_threshold1 = (i,j,k)\n                # print(best_score,best_threshold)\nprint()\ndisplay(best_threshold1,best_score)\n```\n\n\n\n\nit is a slowly idea . then i will try parallel to accelerate this code or try 18 'for' code (It's also not very smart method.)",
      "votes": null
    },
    {
      "id": "2272134",
      "postDate": "05/24/2023 10:14:01",
      "content": "<p>what happened to your CV score when you tuned each threshold for Q1-Q18?</p>",
      "rawMarkdown": "what happened to your CV score when you tuned each threshold for Q1-Q18?",
      "votes": null
    },
    {
      "id": "2272171",
      "postDate": "05/24/2023 10:36:00",
      "content": "<p>The best threshold for CV and the best for LB were slightly different.</p>",
      "rawMarkdown": "The best threshold for CV and the best for LB were slightly different.",
      "votes": null
    },
    {
      "id": "2304797",
      "postDate": "06/16/2023 08:43:11",
      "content": "<p>It seems to work, thank you for sharing.</p>",
      "rawMarkdown": "It seems to work, thank you for sharing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2270698,
      "author_name": "yousefcheng",
      "author_url": "",
      "post_date": "05/23/2023 10:23:12",
      "content": "<p>good idea. i also think this idea. but i only try 3 group 'for' code ,it largely waste time. how do you try q1 to q18 ? use for and other?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2270824,
          "author_name": "masa0take3",
          "author_url": "",
          "post_date": "05/23/2023 13:02:51",
          "content": "<p>I created a list containing 18 thresholds, and executed it with 'for' code.<br>\nI changed these threshold little by little and repeated submit.<br>\nIt's not very smart method.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2270844,
              "author_name": "yousefcheng",
              "author_url": "",
              "post_date": "05/23/2023 13:10:26",
              "content": "<p>thanks, i want to share my idea with you. </p>\n<pre><code>scores = []; thresholds = []\nbest_score = 0; best_threshold1 = 0\n\nfor i in np.arange(0.60,0.65,0.005):\n    for j in np.arange(0.60,0.65,0.005):\n        for k in np.arange(0.60,0.65,0.005):\n            # print(f'{i}{j}{k}',end='')\n            preds=pd.concat([(oof[[0,1,2]]&gt;i).astype('int'),(oof[[3,4,5,6,7,8,9,10,11,12]]&gt;j).astype('int'),(oof[[13,14,15,16,17]]&gt;k).astype('int')],axis=1)\n\n            m = f1_score(true.to_pandas().values.reshape((-1)), preds.to_pandas().values.reshape((-1)), average='macro')   \n            scores.append(m)\n            thresholds.append(threshold)\n            if m&gt;best_score:\n                best_score = m\n\n                best_threshold1 = (i,j,k)\n                # print(best_score,best_threshold)\nprint()\ndisplay(best_threshold1,best_score)\n</code></pre>\n<p>it is a slowly idea . then i will try parallel to accelerate this code or try 18 'for' code (It's also not very smart method.)</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2304797,
                  "author_name": "gojay001",
                  "author_url": "",
                  "post_date": "06/16/2023 08:43:11",
                  "content": "<p>It seems to work, thank you for sharing.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2270701,
      "author_name": "ayaanjang",
      "author_url": "",
      "post_date": "05/23/2023 10:26:04",
      "content": "<p>I think so, too.</p>\n<p>How did you decide on the threshold from Q1 to Q18 individually?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2270833,
          "author_name": "masa0take3",
          "author_url": "",
          "post_date": "05/23/2023 13:05:48",
          "content": "<p>I changed thresholds little by little and repeated submissions.<br>\nI don't know of an efficient way, sorry.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2272134,
      "author_name": "ngocuong",
      "author_url": "",
      "post_date": "05/24/2023 10:14:01",
      "content": "<p>what happened to your CV score when you tuned each threshold for Q1-Q18?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2272171,
          "author_name": "masa0take3",
          "author_url": "",
          "post_date": "05/24/2023 10:36:00",
          "content": "<p>The best threshold for CV and the best for LB were slightly different.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2270692": "Many notebooks set the same threshold for all questions, but I don't think it is the best.\nWhen I set the thresholds for Q1 to Q18 individually, the f1-score increased by about 0.003(LB).\nBut CV doesn't do the same, so it may be overfitting LB.\nWhat would you do?",
    "2270698": "good idea. i also think this idea. but i only try 3 group 'for' code ,it largely waste time. how do you try q1 to q18 ? use for and other?",
    "2270701": "I think so, too.\n\nHow did you decide on the threshold from Q1 to Q18 individually?",
    "2270824": "I created a list containing 18 thresholds, and executed it with 'for' code.\nI changed these threshold little by little and repeated submit.\nIt's not very smart method.",
    "2270833": "I changed thresholds little by little and repeated submissions.\nI don't know of an efficient way, sorry.",
    "2270844": "thanks, i want to share my idea with you. \n\n\n```\nscores = []; thresholds = []\nbest_score = 0; best_threshold1 = 0\n\nfor i in np.arange(0.60,0.65,0.005):\n    for j in np.arange(0.60,0.65,0.005):\n        for k in np.arange(0.60,0.65,0.005):\n            # print(f'{i}{j}{k}',end='')\n            preds=pd.concat([(oof[[0,1,2]]>i).astype('int'),(oof[[3,4,5,6,7,8,9,10,11,12]]>j).astype('int'),(oof[[13,14,15,16,17]]>k).astype('int')],axis=1)\n\n            m = f1_score(true.to_pandas().values.reshape((-1)), preds.to_pandas().values.reshape((-1)), average='macro')   \n            scores.append(m)\n            thresholds.append(threshold)\n            if m>best_score:\n                best_score = m\n                \n                best_threshold1 = (i,j,k)\n                # print(best_score,best_threshold)\nprint()\ndisplay(best_threshold1,best_score)\n```\n\n\n\n\nit is a slowly idea . then i will try parallel to accelerate this code or try 18 'for' code (It's also not very smart method.)",
    "2272134": "what happened to your CV score when you tuned each threshold for Q1-Q18?",
    "2272171": "The best threshold for CV and the best for LB were slightly different.",
    "2304797": "It seems to work, thank you for sharing."
  },
  "source": "meta"
}