{
  "id": 61164,
  "title": "Single Model Performance ",
  "url": "/competitions/freesound-audio-tagging/discussion/61164",
  "author_name": "",
  "post_date": "2018-07-15T05:49:34.377923700Z",
  "votes": 1,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Hello,</p>\n\n<p>Starting a thread to accumulate the single model performances we have around here - to assess the amount of improvements that is still possible.</p>\n\n<p>For example my single model with nearly zero pre-processing (only silence removal + chunking) scores: \nPublic LB: 0.921\nPrivate LB: 0.915\nNumber of Parameters: Less than 1.5M</p>\n\n<p>Same model with silence removal, mixup and chunking scores:\nPublic LB: 0.925\nPrivate LB: 0.907</p>\n\n<p>** EDIT**\nUpdated my current single model performance in both Private and Public LB.\n** END **</p>\n\n<p>Regards,\nGyat</p>",
  "messages": [
    {
      "id": "357031",
      "postDate": "07/15/2018 05:49:34",
      "content": "<p>Hello,</p>\n\n<p>Starting a thread to accumulate the single model performances we have around here - to assess the amount of improvements that is still possible.</p>\n\n<p>For example my single model with nearly zero pre-processing (only silence removal + chunking) scores: \nPublic LB: 0.921\nPrivate LB: 0.915\nNumber of Parameters: Less than 1.5M</p>\n\n<p>Same model with silence removal, mixup and chunking scores:\nPublic LB: 0.925\nPrivate LB: 0.907</p>\n\n<p>** EDIT**\nUpdated my current single model performance in both Private and Public LB.\n** END **</p>\n\n<p>Regards,\nGyat</p>",
      "rawMarkdown": "Hello,\n\nStarting a thread to accumulate the single model performances we have around here - to assess the amount of improvements that is still possible.\n\nFor example my single model with nearly zero pre-processing (only silence removal + chunking) scores: \nPublic LB: 0.921\nPrivate LB: 0.915\nNumber of Parameters: Less than 1.5M\n\nSame model with silence removal, mixup and chunking scores:\nPublic LB: 0.925\nPrivate LB: 0.907\n\n** EDIT**\nUpdated my current single model performance in both Private and Public LB.\n** END **\n\nRegards,\nGyat",
      "votes": null
    },
    {
      "id": "359896",
      "postDate": "07/21/2018 02:46:43",
      "content": "<p>Hi @Gyat,</p>\n\n<p>Sharing score tendency with training dataset:</p>\n\n<ol>\n<li>Trained by focusing on manually verified samples: around 0.92</li>\n<li>Trained for all samples: around 0.93</li>\n<li>Trained with special bias: around 0.94</li>\n</ol>\n\n<p>All the results are with ensemble of single model fold 5 to 10.\n'special bias' above is selecting subset of dataset or re-labeling by using other high performance multiple-model-ensemble result.</p>\n\n<p>Combination of 1 to 3 (= multiple model ensemble) could show higher performance, but I cannot achieve by a single model so far.</p>",
      "rawMarkdown": "Hi @Gyat,\n\nSharing score tendency with training dataset:\n\n1. Trained by focusing on manually verified samples: around 0.92\n2. Trained for all samples: around 0.93\n3. Trained with special bias: around 0.94\n\nAll the results are with ensemble of single model fold 5 to 10.\n'special bias' above is selecting subset of dataset or re-labeling by using other high performance multiple-model-ensemble result.\n\nCombination of 1 to 3 (= multiple model ensemble) could show higher performance, but I cannot achieve by a single model so far.",
      "votes": null
    },
    {
      "id": "360014",
      "postDate": "07/21/2018 09:23:57",
      "content": "<p>This is great! One question. Your 3rd model is using Pseudo Labeling, right?</p>",
      "rawMarkdown": "This is great! One question. Your 3rd model is using Pseudo Labeling, right?",
      "votes": null
    },
    {
      "id": "360098",
      "postDate": "07/21/2018 14:29:09",
      "content": "<p>Hi, no, it's not pseudo labeling. Pseudo labeling from test set is NOT allowed in this competition. It's any of:</p>\n\n<ul>\n<li>Model that learned from sub set of original train set that is selected by former trained model. Or,</li>\n<li>Re-labeled by former trained model; Labels of lower-prediction-prob samples are overwritten (re-labeled) by former trained model's probs.</li>\n</ul>",
      "rawMarkdown": "Hi, no, it's not pseudo labeling. Pseudo labeling from test set is NOT allowed in this competition. It's any of:\n\n- Model that learned from sub set of original train set that is selected by former trained model. Or,\n- Re-labeled by former trained model; Labels of lower-prediction-prob samples are overwritten (re-labeled) by former trained model's probs.",
      "votes": null
    },
    {
      "id": "360770",
      "postDate": "07/23/2018 07:52:05",
      "content": "<p>Hi all,\nThanks for sharing your results. My single model performance is 0.9 (on 10 folds). So i have to work more. As I can see from daisukelab's and our results that ensambling gives 0.02...0.03 more.  Did you use some augmentations or this performance on the raw data?</p>",
      "rawMarkdown": "Hi all,\nThanks for sharing your results. My single model performance is 0.9 (on 10 folds). So i have to work more. As I can see from daisukelab's and our results that ensambling gives 0.02...0.03 more.  Did you use some augmentations or this performance on the raw data?",
      "votes": null
    },
    {
      "id": "360783",
      "postDate": "07/23/2018 08:12:25",
      "content": "<p>I am not sure about <a href=\"/daisukelab\">@daisukelab</a>, but I think my features are a little different from what all has being discussed here. Model is inspired from VGGNET-like architecture. My score is on 5 folds.\nOne unusual thing I'd mention here is that, my model seems to do fine WITHOUT any data augmentations like Chopping into fixed sizes and padding, other augmentations using ImageDataGenerator. Even though I have used it, the marginal benefit is negligible: + 0.002.\nMixup augmentation has been slightly helpful though.</p>",
      "rawMarkdown": "I am not sure about @daisukelab, but I think my features are a little different from what all has being discussed here. Model is inspired from VGGNET-like architecture. My score is on 5 folds.\nOne unusual thing I'd mention here is that, my model seems to do fine WITHOUT any data augmentations like Chopping into fixed sizes and padding, other augmentations using ImageDataGenerator. Even though I have used it, the marginal benefit is negligible: + 0.002.\nMixup augmentation has been slightly helpful though.",
      "votes": null
    },
    {
      "id": "360784",
      "postDate": "07/23/2018 08:14:11",
      "content": "<p>Your third model, the one which scores ~ 0.93, how many parameters does this one have? </p>",
      "rawMarkdown": "Your third model, the one which scores ~ 0.93, how many parameters does this one have?",
      "votes": null
    },
    {
      "id": "360802",
      "postDate": "07/23/2018 09:09:06",
      "content": "<p>It is intrested. I have splited files into fixed size chunks. The architecture was inspired by ResNet. And my model performance is lower.  I think the problem is in chunks with silence. I have to rethink my approach. Thanks!</p>",
      "rawMarkdown": "It is intrested. I have splited files into fixed size chunks. The architecture was inspired by ResNet. And my model performance is lower.  I think the problem is in chunks with silence. I have to rethink my approach. Thanks!",
      "votes": null
    },
    {
      "id": "360889",
      "postDate": "07/23/2018 12:41:23",
      "content": "<p>About 6M, how's your VGGNet like? It sounds great you don't need to augment so much.</p>",
      "rawMarkdown": "About 6M, how's your VGGNet like? It sounds great you don't need to augment so much.",
      "votes": null
    },
    {
      "id": "360891",
      "postDate": "07/23/2018 12:50:31",
      "content": "<p>Hi Oleh, augmentation matters and results in +0.5 or more, comparing to less augmentation training in my environment. But your model could be good enough. You might also consider followings:</p>\n\n<ul>\n<li>Preprocessing: Using 44.1kHz as is, changing split time length and more.</li>\n<li>Dataset: Using manually verified, forming subset of training set or anything.</li>\n</ul>",
      "rawMarkdown": "Hi Oleh, augmentation matters and results in +0.5 or more, comparing to less augmentation training in my environment. But your model could be good enough. You might also consider followings:\n\n- Preprocessing: Using 44.1kHz as is, changing split time length and more.\n- Dataset: Using manually verified, forming subset of training set or anything.",
      "votes": null
    },
    {
      "id": "360892",
      "postDate": "07/23/2018 12:50:45",
      "content": "<p>My model has only about 1.5M parameters. \nBut then again the current LB only considers 19% of the training set. The picture may be very different and my model might take the downward plunge for the rest 81% :-/\nHow well your validation loss/accuracy relates to the MAP@3 score in the LB? That is, around what value of val loss/accuracy got you at 0.93 in the LB?</p>",
      "rawMarkdown": "My model has only about 1.5M parameters. \nBut then again the current LB only considers 19% of the training set. The picture may be very different and my model might take the downward plunge for the rest 81% :-/\nHow well your validation loss/accuracy relates to the MAP@3 score in the LB? That is, around what value of val loss/accuracy got you at 0.93 in the LB?",
      "votes": null
    },
    {
      "id": "360907",
      "postDate": "07/23/2018 13:32:58",
      "content": "<p>Hi, daisukelab, thanks for your reply,  I have used 44.1kHz but did not change split time length. I will try to play around split time. Also, I am going to try manually verified audio only. Because when I experimented with silence removal from the audio I have faced with this guy 93198de4.wav. Labeled as Snare drum, but I hear 3 different kinds of drums in this audio (including the intersection with Hi-hat class). So I will try manually verified samples in the next experiment. Thanks</p>",
      "rawMarkdown": "Hi, daisukelab, thanks for your reply,  I have used 44.1kHz but did not change split time length. I will try to play around split time. Also, I am going to try manually verified audio only. Because when I experimented with silence removal from the audio I have faced with this guy 93198de4.wav. Labeled as Snare drum, but I hear 3 different kinds of drums in this audio (including the intersection with Hi-hat class). So I will try manually verified samples in the next experiment. Thanks",
      "votes": null
    },
    {
      "id": "360934",
      "postDate": "07/23/2018 14:39:26",
      "content": "<p>Hi Gyat, thank you for reminding me of the current LB scoring.</p>\n\n<ul>\n<li>Attempt E9 (9 folds) typical best - val_loss: 0.3577 - val_acc: 0.9189</li>\n<li>Attempt M9P615 (10 folds) typical best - val_loss: 0.3713 - val_acc: 0.9180</li>\n</ul>\n\n<p>(leaving local attempt name for my reference...)</p>\n\n<p>These two show over 0.93 in the LB for example. My understanding was the LB score 0.93 is achieved by  ensemble of many 0.918s, but it could also be possible that these result are just high in the current LB...</p>",
      "rawMarkdown": "Hi Gyat, thank you for reminding me of the current LB scoring.\n\n- Attempt E9 (9 folds) typical best - val_loss: 0.3577 - val_acc: 0.9189\n- Attempt M9P615 (10 folds) typical best - val_loss: 0.3713 - val_acc: 0.9180\n\n(leaving local attempt name for my reference...)\n\nThese two show over 0.93 in the LB for example. My understanding was the LB score 0.93 is achieved by  ensemble of many 0.918s, but it could also be possible that these result are just high in the current LB...",
      "votes": null
    },
    {
      "id": "360973",
      "postDate": "07/23/2018 15:44:37",
      "content": "<p>Oh boy! My val-acc is wayyyy lower than yours. This is a new cause for worry. Seems like my model might fail for the rest 81% data.</p>\n\n<p>Is your validation set based on the noisy weakly labeled data? Or is it only on the manually verified one?\nAlso, could you share your process of creating the train/validation splits?</p>",
      "rawMarkdown": "Oh boy! My val-acc is wayyyy lower than yours. This is a new cause for worry. Seems like my model might fail for the rest 81% data.\n\nIs your validation set based on the noisy weakly labeled data? Or is it only on the manually verified one?\nAlso, could you share your process of creating the train/validation splits?",
      "votes": null
    },
    {
      "id": "361142",
      "postDate": "07/23/2018 23:21:10",
      "content": "<p>Hey I'm kind of old boy ;)</p>\n\n<ul>\n<li>These val_loss/acc are based on all train set samples, all including weakly labeled ones.</li>\n<li>test_size = 0.2</li>\n</ul>",
      "rawMarkdown": "Hey I'm kind of old boy ;)\n\n- These val_loss/acc are based on all train set samples, all including weakly labeled ones.\n- test_size = 0.2",
      "votes": null
    },
    {
      "id": "361260",
      "postDate": "07/24/2018 05:39:33",
      "content": "<p>Worried :-/</p>",
      "rawMarkdown": "Worried :-/",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 359896,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "07/21/2018 02:46:43",
      "content": "<p>Hi @Gyat,</p>\n\n<p>Sharing score tendency with training dataset:</p>\n\n<ol>\n<li>Trained by focusing on manually verified samples: around 0.92</li>\n<li>Trained for all samples: around 0.93</li>\n<li>Trained with special bias: around 0.94</li>\n</ol>\n\n<p>All the results are with ensemble of single model fold 5 to 10.\n'special bias' above is selecting subset of dataset or re-labeling by using other high performance multiple-model-ensemble result.</p>\n\n<p>Combination of 1 to 3 (= multiple model ensemble) could show higher performance, but I cannot achieve by a single model so far.</p>",
      "votes": null,
      "replies": [
        {
          "id": 360014,
          "author_name": "gyat2017",
          "author_url": "",
          "post_date": "07/21/2018 09:23:57",
          "content": "<p>This is great! One question. Your 3rd model is using Pseudo Labeling, right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 360098,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "07/21/2018 14:29:09",
          "content": "<p>Hi, no, it's not pseudo labeling. Pseudo labeling from test set is NOT allowed in this competition. It's any of:</p>\n\n<ul>\n<li>Model that learned from sub set of original train set that is selected by former trained model. Or,</li>\n<li>Re-labeled by former trained model; Labels of lower-prediction-prob samples are overwritten (re-labeled) by former trained model's probs.</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 360784,
          "author_name": "gyat2017",
          "author_url": "",
          "post_date": "07/23/2018 08:14:11",
          "content": "<p>Your third model, the one which scores ~ 0.93, how many parameters does this one have? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 360889,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "07/23/2018 12:41:23",
          "content": "<p>About 6M, how's your VGGNet like? It sounds great you don't need to augment so much.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 360892,
          "author_name": "gyat2017",
          "author_url": "",
          "post_date": "07/23/2018 12:50:45",
          "content": "<p>My model has only about 1.5M parameters. \nBut then again the current LB only considers 19% of the training set. The picture may be very different and my model might take the downward plunge for the rest 81% :-/\nHow well your validation loss/accuracy relates to the MAP@3 score in the LB? That is, around what value of val loss/accuracy got you at 0.93 in the LB?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 360934,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "07/23/2018 14:39:26",
          "content": "<p>Hi Gyat, thank you for reminding me of the current LB scoring.</p>\n\n<ul>\n<li>Attempt E9 (9 folds) typical best - val_loss: 0.3577 - val_acc: 0.9189</li>\n<li>Attempt M9P615 (10 folds) typical best - val_loss: 0.3713 - val_acc: 0.9180</li>\n</ul>\n\n<p>(leaving local attempt name for my reference...)</p>\n\n<p>These two show over 0.93 in the LB for example. My understanding was the LB score 0.93 is achieved by  ensemble of many 0.918s, but it could also be possible that these result are just high in the current LB...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 360973,
          "author_name": "gyat2017",
          "author_url": "",
          "post_date": "07/23/2018 15:44:37",
          "content": "<p>Oh boy! My val-acc is wayyyy lower than yours. This is a new cause for worry. Seems like my model might fail for the rest 81% data.</p>\n\n<p>Is your validation set based on the noisy weakly labeled data? Or is it only on the manually verified one?\nAlso, could you share your process of creating the train/validation splits?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 361142,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "07/23/2018 23:21:10",
          "content": "<p>Hey I'm kind of old boy ;)</p>\n\n<ul>\n<li>These val_loss/acc are based on all train set samples, all including weakly labeled ones.</li>\n<li>test_size = 0.2</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 361260,
          "author_name": "gyat2017",
          "author_url": "",
          "post_date": "07/24/2018 05:39:33",
          "content": "<p>Worried :-/</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 360770,
      "author_name": "bodilowsky",
      "author_url": "",
      "post_date": "07/23/2018 07:52:05",
      "content": "<p>Hi all,\nThanks for sharing your results. My single model performance is 0.9 (on 10 folds). So i have to work more. As I can see from daisukelab's and our results that ensambling gives 0.02...0.03 more.  Did you use some augmentations or this performance on the raw data?</p>",
      "votes": null,
      "replies": [
        {
          "id": 360783,
          "author_name": "gyat2017",
          "author_url": "",
          "post_date": "07/23/2018 08:12:25",
          "content": "<p>I am not sure about <a href=\"/daisukelab\">@daisukelab</a>, but I think my features are a little different from what all has being discussed here. Model is inspired from VGGNET-like architecture. My score is on 5 folds.\nOne unusual thing I'd mention here is that, my model seems to do fine WITHOUT any data augmentations like Chopping into fixed sizes and padding, other augmentations using ImageDataGenerator. Even though I have used it, the marginal benefit is negligible: + 0.002.\nMixup augmentation has been slightly helpful though.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 360802,
          "author_name": "bodilowsky",
          "author_url": "",
          "post_date": "07/23/2018 09:09:06",
          "content": "<p>It is intrested. I have splited files into fixed size chunks. The architecture was inspired by ResNet. And my model performance is lower.  I think the problem is in chunks with silence. I have to rethink my approach. Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 360891,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "07/23/2018 12:50:31",
          "content": "<p>Hi Oleh, augmentation matters and results in +0.5 or more, comparing to less augmentation training in my environment. But your model could be good enough. You might also consider followings:</p>\n\n<ul>\n<li>Preprocessing: Using 44.1kHz as is, changing split time length and more.</li>\n<li>Dataset: Using manually verified, forming subset of training set or anything.</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 360907,
          "author_name": "bodilowsky",
          "author_url": "",
          "post_date": "07/23/2018 13:32:58",
          "content": "<p>Hi, daisukelab, thanks for your reply,  I have used 44.1kHz but did not change split time length. I will try to play around split time. Also, I am going to try manually verified audio only. Because when I experimented with silence removal from the audio I have faced with this guy 93198de4.wav. Labeled as Snare drum, but I hear 3 different kinds of drums in this audio (including the intersection with Hi-hat class). So I will try manually verified samples in the next experiment. Thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "357031": "Hello,\n\nStarting a thread to accumulate the single model performances we have around here - to assess the amount of improvements that is still possible.\n\nFor example my single model with nearly zero pre-processing (only silence removal + chunking) scores: \nPublic LB: 0.921\nPrivate LB: 0.915\nNumber of Parameters: Less than 1.5M\n\nSame model with silence removal, mixup and chunking scores:\nPublic LB: 0.925\nPrivate LB: 0.907\n\n** EDIT**\nUpdated my current single model performance in both Private and Public LB.\n** END **\n\nRegards,\nGyat",
    "359896": "Hi @Gyat,\n\nSharing score tendency with training dataset:\n\n1. Trained by focusing on manually verified samples: around 0.92\n2. Trained for all samples: around 0.93\n3. Trained with special bias: around 0.94\n\nAll the results are with ensemble of single model fold 5 to 10.\n'special bias' above is selecting subset of dataset or re-labeling by using other high performance multiple-model-ensemble result.\n\nCombination of 1 to 3 (= multiple model ensemble) could show higher performance, but I cannot achieve by a single model so far.",
    "360014": "This is great! One question. Your 3rd model is using Pseudo Labeling, right?",
    "360098": "Hi, no, it's not pseudo labeling. Pseudo labeling from test set is NOT allowed in this competition. It's any of:\n\n- Model that learned from sub set of original train set that is selected by former trained model. Or,\n- Re-labeled by former trained model; Labels of lower-prediction-prob samples are overwritten (re-labeled) by former trained model's probs.",
    "360770": "Hi all,\nThanks for sharing your results. My single model performance is 0.9 (on 10 folds). So i have to work more. As I can see from daisukelab's and our results that ensambling gives 0.02...0.03 more.  Did you use some augmentations or this performance on the raw data?",
    "360783": "I am not sure about @daisukelab, but I think my features are a little different from what all has being discussed here. Model is inspired from VGGNET-like architecture. My score is on 5 folds.\nOne unusual thing I'd mention here is that, my model seems to do fine WITHOUT any data augmentations like Chopping into fixed sizes and padding, other augmentations using ImageDataGenerator. Even though I have used it, the marginal benefit is negligible: + 0.002.\nMixup augmentation has been slightly helpful though.",
    "360784": "Your third model, the one which scores ~ 0.93, how many parameters does this one have?",
    "360802": "It is intrested. I have splited files into fixed size chunks. The architecture was inspired by ResNet. And my model performance is lower.  I think the problem is in chunks with silence. I have to rethink my approach. Thanks!",
    "360889": "About 6M, how's your VGGNet like? It sounds great you don't need to augment so much.",
    "360891": "Hi Oleh, augmentation matters and results in +0.5 or more, comparing to less augmentation training in my environment. But your model could be good enough. You might also consider followings:\n\n- Preprocessing: Using 44.1kHz as is, changing split time length and more.\n- Dataset: Using manually verified, forming subset of training set or anything.",
    "360892": "My model has only about 1.5M parameters. \nBut then again the current LB only considers 19% of the training set. The picture may be very different and my model might take the downward plunge for the rest 81% :-/\nHow well your validation loss/accuracy relates to the MAP@3 score in the LB? That is, around what value of val loss/accuracy got you at 0.93 in the LB?",
    "360907": "Hi, daisukelab, thanks for your reply,  I have used 44.1kHz but did not change split time length. I will try to play around split time. Also, I am going to try manually verified audio only. Because when I experimented with silence removal from the audio I have faced with this guy 93198de4.wav. Labeled as Snare drum, but I hear 3 different kinds of drums in this audio (including the intersection with Hi-hat class). So I will try manually verified samples in the next experiment. Thanks",
    "360934": "Hi Gyat, thank you for reminding me of the current LB scoring.\n\n- Attempt E9 (9 folds) typical best - val_loss: 0.3577 - val_acc: 0.9189\n- Attempt M9P615 (10 folds) typical best - val_loss: 0.3713 - val_acc: 0.9180\n\n(leaving local attempt name for my reference...)\n\nThese two show over 0.93 in the LB for example. My understanding was the LB score 0.93 is achieved by  ensemble of many 0.918s, but it could also be possible that these result are just high in the current LB...",
    "360973": "Oh boy! My val-acc is wayyyy lower than yours. This is a new cause for worry. Seems like my model might fail for the rest 81% data.\n\nIs your validation set based on the noisy weakly labeled data? Or is it only on the manually verified one?\nAlso, could you share your process of creating the train/validation splits?",
    "361142": "Hey I'm kind of old boy ;)\n\n- These val_loss/acc are based on all train set samples, all including weakly labeled ones.\n- test_size = 0.2",
    "361260": "Worried :-/"
  },
  "source": "meta"
}