{
  "id": 181141,
  "title": "Good CV but low lb score",
  "url": "/competitions/birdsong-recognition/discussion/181141",
  "author_name": "",
  "post_date": "2020-09-07T18:08:46.264084400Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I got a cv f1 around 0.65 but my lb score was around 0.44 Can anyone explain why this happens??I used effnet tpu-train and inference as my baseline model.And also while submitting inference, I submitted only in cpu not gpu cuz it has thrown notebook timeout. Has anybody faced this error?</p>",
  "messages": [
    {
      "id": "1001993",
      "postDate": "09/07/2020 18:08:46",
      "content": "<p>I got a cv f1 around 0.65 but my lb score was around 0.44 Can anyone explain why this happens??I used effnet tpu-train and inference as my baseline model.And also while submitting inference, I submitted only in cpu not gpu cuz it has thrown notebook timeout. Has anybody faced this error?</p>",
      "rawMarkdown": "I got a cv f1 around 0.65 but my lb score was around 0.44 Can anyone explain why this happens??I used effnet tpu-train and inference as my baseline model.And also while submitting inference, I submitted only in cpu not gpu cuz it has thrown notebook timeout. Has anybody faced this error?",
      "votes": null
    },
    {
      "id": "1002058",
      "postDate": "09/07/2020 19:28:03",
      "content": "<p>I have faced the error and I'm still facing it. My CV-Score is in the same region as yours ~0.60 but my LB-Score is very low and varies from 0.009 to 0.534 even with using <code>early_stopping</code> so it is not an overfitting issue.</p>\n<p>When it comes to <code>notebook timeout</code> you should cache your soundfiles, my submission takes around 10-15 minutes on GPU.</p>\n<p>I implemented it in my dataset, I show you how:<br>\nFirst define a dict in your dataset:</p>\n<p><code>self.cache = {'filename': None, 'soundfile': None}</code></p>\n<p>Then when loading the audio files you can just check if its already in cache otherwise you save it as a variable in a dict:</p>\n<pre><code>if self.cache['filename'] == filename:\n    soundfile = self.cache['soundfile']\nelse:\n    soundfile = AudioSegment.from_mp3(path)\n    self.cache['filename'] = filename\n    self.cache['soundfile'] = soundfile\n</code></pre>",
      "rawMarkdown": "I have faced the error and I'm still facing it. My CV-Score is in the same region as yours ~0.60 but my LB-Score is very low and varies from 0.009 to 0.534 even with using `early_stopping` so it is not an overfitting issue.\n\nWhen it comes to `notebook timeout` you should cache your soundfiles, my submission takes around 10-15 minutes on GPU.\n\nI implemented it in my dataset, I show you how:\nFirst define a dict in your dataset:\n\n`self.cache = {'filename': None, 'soundfile': None}`\n\nThen when loading the audio files you can just check if its already in cache otherwise you save it as a variable in a dict:\n\n```\nif self.cache['filename'] == filename:\n    soundfile = self.cache['soundfile']\nelse:\n    soundfile = AudioSegment.from_mp3(path)\n    self.cache['filename'] = filename\n    self.cache['soundfile'] = soundfile\n```",
      "votes": null
    },
    {
      "id": "1002263",
      "postDate": "09/08/2020 01:58:51",
      "content": "<p>Thanks 4 ur comment</p>",
      "rawMarkdown": "Thanks 4 ur comment",
      "votes": null
    },
    {
      "id": "1002422",
      "postDate": "09/08/2020 05:43:05",
      "content": "<p>This means that your model has \"learnt well\"  on train set like  conditions.<br>\nBut the test set is very noisy and soundscape based and your model should 'adapt' and predict on that. This is problem that the organisers want to solve.</p>",
      "rawMarkdown": "This means that your model has \"learnt well\"  on train set like  conditions.\nBut the test set is very noisy and soundscape based and your model should 'adapt' and predict on that. This is problem that the organisers want to solve.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1002058,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "09/07/2020 19:28:03",
      "content": "<p>I have faced the error and I'm still facing it. My CV-Score is in the same region as yours ~0.60 but my LB-Score is very low and varies from 0.009 to 0.534 even with using <code>early_stopping</code> so it is not an overfitting issue.</p>\n<p>When it comes to <code>notebook timeout</code> you should cache your soundfiles, my submission takes around 10-15 minutes on GPU.</p>\n<p>I implemented it in my dataset, I show you how:<br>\nFirst define a dict in your dataset:</p>\n<p><code>self.cache = {'filename': None, 'soundfile': None}</code></p>\n<p>Then when loading the audio files you can just check if its already in cache otherwise you save it as a variable in a dict:</p>\n<pre><code>if self.cache['filename'] == filename:\n    soundfile = self.cache['soundfile']\nelse:\n    soundfile = AudioSegment.from_mp3(path)\n    self.cache['filename'] = filename\n    self.cache['soundfile'] = soundfile\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1002263,
          "author_name": "vineeth1999",
          "author_url": "",
          "post_date": "09/08/2020 01:58:51",
          "content": "<p>Thanks 4 ur comment</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1002422,
      "author_name": "watzisname",
      "author_url": "",
      "post_date": "09/08/2020 05:43:05",
      "content": "<p>This means that your model has \"learnt well\"  on train set like  conditions.<br>\nBut the test set is very noisy and soundscape based and your model should 'adapt' and predict on that. This is problem that the organisers want to solve.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1001993": "I got a cv f1 around 0.65 but my lb score was around 0.44 Can anyone explain why this happens??I used effnet tpu-train and inference as my baseline model.And also while submitting inference, I submitted only in cpu not gpu cuz it has thrown notebook timeout. Has anybody faced this error?",
    "1002058": "I have faced the error and I'm still facing it. My CV-Score is in the same region as yours ~0.60 but my LB-Score is very low and varies from 0.009 to 0.534 even with using `early_stopping` so it is not an overfitting issue.\n\nWhen it comes to `notebook timeout` you should cache your soundfiles, my submission takes around 10-15 minutes on GPU.\n\nI implemented it in my dataset, I show you how:\nFirst define a dict in your dataset:\n\n`self.cache = {'filename': None, 'soundfile': None}`\n\nThen when loading the audio files you can just check if its already in cache otherwise you save it as a variable in a dict:\n\n```\nif self.cache['filename'] == filename:\n    soundfile = self.cache['soundfile']\nelse:\n    soundfile = AudioSegment.from_mp3(path)\n    self.cache['filename'] = filename\n    self.cache['soundfile'] = soundfile\n```",
    "1002263": "Thanks 4 ur comment",
    "1002422": "This means that your model has \"learnt well\"  on train set like  conditions.\nBut the test set is very noisy and soundscape based and your model should 'adapt' and predict on that. This is problem that the organisers want to solve."
  },
  "source": "meta"
}