{
  "id": 492518,
  "title": "Model sanity check: Reverse your predictions",
  "url": "/competitions/birdclef-2024/discussion/492518",
  "author_name": "",
  "post_date": "2024-04-09T23:23:06.409446300Z",
  "votes": 3,
  "comment_count": 7,
  "views": 0,
  "content": "<p>So - I've made a few submissions at this point - my best being .56 on the LB</p>\n<p>I'd been making some tweaks to my model that I'd hope would improve performance but have not (scoring around .55).  This made me question if my models were performing at all - wanted to try to verify that…</p>\n<p>Obviously - anything significantly above .50 should indicate your model is working - but I was concerned maybe there was some kind of scoring oddity.</p>\n<p>I decided to do a simple test - and reverse the predicted columns - so the classes were all incorrect.</p>\n<p>predictions = predictions[::-1]</p>\n<p>I was happy to see my score drop to .49 on the LB.</p>\n<p>Just thought I'd share in case this is useful to anyone.</p>",
  "messages": [
    {
      "id": "2744446",
      "postDate": "04/09/2024 23:23:06",
      "content": "<p>So - I've made a few submissions at this point - my best being .56 on the LB</p>\n<p>I'd been making some tweaks to my model that I'd hope would improve performance but have not (scoring around .55).  This made me question if my models were performing at all - wanted to try to verify that…</p>\n<p>Obviously - anything significantly above .50 should indicate your model is working - but I was concerned maybe there was some kind of scoring oddity.</p>\n<p>I decided to do a simple test - and reverse the predicted columns - so the classes were all incorrect.</p>\n<p>predictions = predictions[::-1]</p>\n<p>I was happy to see my score drop to .49 on the LB.</p>\n<p>Just thought I'd share in case this is useful to anyone.</p>",
      "rawMarkdown": "So - I've made a few submissions at this point - my best being .56 on the LB\n\nI'd been making some tweaks to my model that I'd hope would improve performance but have not (scoring around .55).  This made me question if my models were performing at all - wanted to try to verify that...\n\nObviously - anything significantly above .50 should indicate your model is working - but I was concerned maybe there was some kind of scoring oddity.\n\nI decided to do a simple test - and reverse the predicted columns - so the classes were all incorrect.\n\npredictions = predictions[::-1]\n\nI was happy to see my score drop to .49 on the LB.\n\nJust thought I'd share in case this is useful to anyone.",
      "votes": null
    },
    {
      "id": "2744676",
      "postDate": "04/10/2024 04:20:59",
      "content": "<p>My model is still insane, its been 3 days and I m still not able to figure out what's wrong with my model, idk if its the model or the training label encoding I can't get pass 0.5 </p>",
      "rawMarkdown": "My model is still insane, its been 3 days and I m still not able to figure out what's wrong with my model, idk if its the model or the training label encoding I can't get pass 0.5",
      "votes": null
    },
    {
      "id": "2744781",
      "postDate": "04/10/2024 05:56:42",
      "content": "<p>are you doing separate notebooks for train and run?</p>\n<p>I had a bunch of problems getting my model to load in my \"run\" notebook correctly.  It looked like the model was loading - but in one case it was prediction the same values for all classes.  Another time the classes weren't getting mapped right….</p>\n<p>If you're training separately - I'd try doing at least one prediction from your \"train\" notebook - and then do the same prediction in your \"run\" notebook.  Verify the problem isn't with save / loading.</p>",
      "rawMarkdown": "are you doing separate notebooks for train and run?\n\nI had a bunch of problems getting my model to load in my \"run\" notebook correctly.  It looked like the model was loading - but in one case it was prediction the same values for all classes.  Another time the classes weren't getting mapped right....\n\nIf you're training separately - I'd try doing at least one prediction from your \"train\" notebook - and then do the same prediction in your \"run\" notebook.  Verify the problem isn't with save / loading.",
      "votes": null
    },
    {
      "id": "2744912",
      "postDate": "04/10/2024 07:54:26",
      "content": "<p>Thanks for the suggestion, I did try to check them out, I tried using the unlabeled data to see if anything is wrong, I haven't found any so far, would you mind checking this notebook out and see where I am getting it wrong? <a href=\"https://www.kaggle.com/arunsensei/baseline-train-efficientnetb0\" target=\"_blank\">notebook</a>, I would really appreciate some help if you are willing to point out the mistake in that, the traindata I used is small since I am just trying it out but even tho I would atleast except 0.53 or 0.54 to verify if everything's right<br>\nThank you,</p>",
      "rawMarkdown": "Thanks for the suggestion, I did try to check them out, I tried using the unlabeled data to see if anything is wrong, I haven't found any so far, would you mind checking this notebook out and see where I am getting it wrong? [notebook](https://www.kaggle.com/arunsensei/baseline-train-efficientnetb0), I would really appreciate some help if you are willing to point out the mistake in that, the traindata I used is small since I am just trying it out but even tho I would atleast except 0.53 or 0.54 to verify if everything's right\nThank you,",
      "votes": null
    },
    {
      "id": "2745737",
      "postDate": "04/10/2024 19:10:27",
      "content": "<p>I'd be happy to take a look - but didn't look like the link to your notebook came through.</p>",
      "rawMarkdown": "I'd be happy to take a look - but didn't look like the link to your notebook came through.",
      "votes": null
    },
    {
      "id": "2747323",
      "postDate": "04/11/2024 19:54:27",
      "content": "<p>thanks for willing to help, I did provide the <a href=\"https://www.kaggle.com/code/arunsensei/baseline-train-efficientnetb0\" target=\"_blank\">link</a>, but I think I figured out what's I was doing wrong, bad preprocessing I think</p>",
      "rawMarkdown": "thanks for willing to help, I did provide the [link](https://www.kaggle.com/code/arunsensei/baseline-train-efficientnetb0), but I think I figured out what's I was doing wrong, bad preprocessing I think",
      "votes": null
    },
    {
      "id": "2747366",
      "postDate": "04/11/2024 20:29:04",
      "content": "<p>oops! - I see the link now (the underline just didn't stand out for me).  </p>\n<p>so - I -might- see an issue…. </p>\n<p>are you doing anything to assure that the class order matches the column order when you predict?</p>\n<p>specifically - this:<br>\nos.listdir(train_path)</p>\n<p>doesn't necessarily return an alphabetically list of folders.  so - unless you account for that somehow - the order won't necessarily match the alpha-sorted columns</p>\n<p>from <a href=\"https://www.kaggle.com/code/richolson/birdclef-2024-spectrograms-imagenet-train\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-2024-spectrograms-imagenet-train</a></p>\n<pre><code>\n\nclass_labels = {class_name: i  i, class_name  ((os.listdir(image_folder)))}\nnum_classes = (class_labels)\n</code></pre>\n<p>might not be related - but a thought.</p>",
      "rawMarkdown": "oops! - I see the link now (the underline just didn't stand out for me).  \n\nso - I -might- see an issue.... \n\nare you doing anything to assure that the class order matches the column order when you predict?\n\nspecifically - this:\nos.listdir(train_path)\n\ndoesn't necessarily return an alphabetically list of folders.  so - unless you account for that somehow - the order won't necessarily match the alpha-sorted columns\n\nfrom https://www.kaggle.com/code/richolson/birdclef-2024-spectrograms-imagenet-train\n```python\n# Mapping of classes to numerical labels\n# Classes are alphabetically sorted - so we can easily restore class order when we load\nclass_labels = {class_name: i for i, class_name in enumerate(sorted(os.listdir(image_folder)))}\nnum_classes = len(class_labels)\n\n```\n\nmight not be related - but a thought.",
      "votes": null
    },
    {
      "id": "2747569",
      "postDate": "04/12/2024 01:09:59",
      "content": "<p>Thanks for the help. But the actual problem was I was using too little data and it's getting overfit on that, scoring improved after I used more data in train </p>",
      "rawMarkdown": "Thanks for the help. But the actual problem was I was using too little data and it's getting overfit on that, scoring improved after I used more data in train",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2744676,
      "author_name": "arunsensei",
      "author_url": "",
      "post_date": "04/10/2024 04:20:59",
      "content": "<p>My model is still insane, its been 3 days and I m still not able to figure out what's wrong with my model, idk if its the model or the training label encoding I can't get pass 0.5 </p>",
      "votes": null,
      "replies": [
        {
          "id": 2744781,
          "author_name": "richolson",
          "author_url": "",
          "post_date": "04/10/2024 05:56:42",
          "content": "<p>are you doing separate notebooks for train and run?</p>\n<p>I had a bunch of problems getting my model to load in my \"run\" notebook correctly.  It looked like the model was loading - but in one case it was prediction the same values for all classes.  Another time the classes weren't getting mapped right….</p>\n<p>If you're training separately - I'd try doing at least one prediction from your \"train\" notebook - and then do the same prediction in your \"run\" notebook.  Verify the problem isn't with save / loading.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2744912,
              "author_name": "arunsensei",
              "author_url": "",
              "post_date": "04/10/2024 07:54:26",
              "content": "<p>Thanks for the suggestion, I did try to check them out, I tried using the unlabeled data to see if anything is wrong, I haven't found any so far, would you mind checking this notebook out and see where I am getting it wrong? <a href=\"https://www.kaggle.com/arunsensei/baseline-train-efficientnetb0\" target=\"_blank\">notebook</a>, I would really appreciate some help if you are willing to point out the mistake in that, the traindata I used is small since I am just trying it out but even tho I would atleast except 0.53 or 0.54 to verify if everything's right<br>\nThank you,</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2745737,
                  "author_name": "richolson",
                  "author_url": "",
                  "post_date": "04/10/2024 19:10:27",
                  "content": "<p>I'd be happy to take a look - but didn't look like the link to your notebook came through.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2747323,
                      "author_name": "arunsensei",
                      "author_url": "",
                      "post_date": "04/11/2024 19:54:27",
                      "content": "<p>thanks for willing to help, I did provide the <a href=\"https://www.kaggle.com/code/arunsensei/baseline-train-efficientnetb0\" target=\"_blank\">link</a>, but I think I figured out what's I was doing wrong, bad preprocessing I think</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2747366,
                          "author_name": "richolson",
                          "author_url": "",
                          "post_date": "04/11/2024 20:29:04",
                          "content": "<p>oops! - I see the link now (the underline just didn't stand out for me).  </p>\n<p>so - I -might- see an issue…. </p>\n<p>are you doing anything to assure that the class order matches the column order when you predict?</p>\n<p>specifically - this:<br>\nos.listdir(train_path)</p>\n<p>doesn't necessarily return an alphabetically list of folders.  so - unless you account for that somehow - the order won't necessarily match the alpha-sorted columns</p>\n<p>from <a href=\"https://www.kaggle.com/code/richolson/birdclef-2024-spectrograms-imagenet-train\" target=\"_blank\">https://www.kaggle.com/code/richolson/birdclef-2024-spectrograms-imagenet-train</a></p>\n<pre><code>\n\nclass_labels = {class_name: i  i, class_name  ((os.listdir(image_folder)))}\nnum_classes = (class_labels)\n</code></pre>\n<p>might not be related - but a thought.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2747569,
                              "author_name": "arunsensei",
                              "author_url": "",
                              "post_date": "04/12/2024 01:09:59",
                              "content": "<p>Thanks for the help. But the actual problem was I was using too little data and it's getting overfit on that, scoring improved after I used more data in train </p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2744446": "So - I've made a few submissions at this point - my best being .56 on the LB\n\nI'd been making some tweaks to my model that I'd hope would improve performance but have not (scoring around .55).  This made me question if my models were performing at all - wanted to try to verify that...\n\nObviously - anything significantly above .50 should indicate your model is working - but I was concerned maybe there was some kind of scoring oddity.\n\nI decided to do a simple test - and reverse the predicted columns - so the classes were all incorrect.\n\npredictions = predictions[::-1]\n\nI was happy to see my score drop to .49 on the LB.\n\nJust thought I'd share in case this is useful to anyone.",
    "2744676": "My model is still insane, its been 3 days and I m still not able to figure out what's wrong with my model, idk if its the model or the training label encoding I can't get pass 0.5",
    "2744781": "are you doing separate notebooks for train and run?\n\nI had a bunch of problems getting my model to load in my \"run\" notebook correctly.  It looked like the model was loading - but in one case it was prediction the same values for all classes.  Another time the classes weren't getting mapped right....\n\nIf you're training separately - I'd try doing at least one prediction from your \"train\" notebook - and then do the same prediction in your \"run\" notebook.  Verify the problem isn't with save / loading.",
    "2744912": "Thanks for the suggestion, I did try to check them out, I tried using the unlabeled data to see if anything is wrong, I haven't found any so far, would you mind checking this notebook out and see where I am getting it wrong? [notebook](https://www.kaggle.com/arunsensei/baseline-train-efficientnetb0), I would really appreciate some help if you are willing to point out the mistake in that, the traindata I used is small since I am just trying it out but even tho I would atleast except 0.53 or 0.54 to verify if everything's right\nThank you,",
    "2745737": "I'd be happy to take a look - but didn't look like the link to your notebook came through.",
    "2747323": "thanks for willing to help, I did provide the [link](https://www.kaggle.com/code/arunsensei/baseline-train-efficientnetb0), but I think I figured out what's I was doing wrong, bad preprocessing I think",
    "2747366": "oops! - I see the link now (the underline just didn't stand out for me).  \n\nso - I -might- see an issue.... \n\nare you doing anything to assure that the class order matches the column order when you predict?\n\nspecifically - this:\nos.listdir(train_path)\n\ndoesn't necessarily return an alphabetically list of folders.  so - unless you account for that somehow - the order won't necessarily match the alpha-sorted columns\n\nfrom https://www.kaggle.com/code/richolson/birdclef-2024-spectrograms-imagenet-train\n```python\n# Mapping of classes to numerical labels\n# Classes are alphabetically sorted - so we can easily restore class order when we load\nclass_labels = {class_name: i for i, class_name in enumerate(sorted(os.listdir(image_folder)))}\nnum_classes = len(class_labels)\n\n```\n\nmight not be related - but a thought.",
    "2747569": "Thanks for the help. But the actual problem was I was using too little data and it's getting overfit on that, scoring improved after I used more data in train"
  },
  "source": "meta"
}