{
  "id": 128527,
  "title": "Is there any way to evaluate a model at an early stage?",
  "url": "/competitions/bengaliai-cv19/discussion/128527",
  "author_name": "",
  "post_date": "2020-02-01T01:52:16.402110600Z",
  "votes": 3,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi all,\nIs there any way to evaluate a model at an early stage, e.g, 50 epochs? It is recommended in many posts that training 100 epochs makes the model finally converge. But it takes me 33 hours to train 100 epochs.\nIs it convincing if I compare two models' performance if they haven't finally converged?  For instance, if there're two models, one's recall is 0.945 at 50 epochs, the other's is 0.949. Can I say the second's recall is better than the first at 100 epochs?</p>\n\n<p>My local valid recall is 0.944 at 50 epochs.</p>",
  "messages": [
    {
      "id": "734147",
      "postDate": "02/01/2020 01:52:16",
      "content": "<p>Hi all,\nIs there any way to evaluate a model at an early stage, e.g, 50 epochs? It is recommended in many posts that training 100 epochs makes the model finally converge. But it takes me 33 hours to train 100 epochs.\nIs it convincing if I compare two models' performance if they haven't finally converged?  For instance, if there're two models, one's recall is 0.945 at 50 epochs, the other's is 0.949. Can I say the second's recall is better than the first at 100 epochs?</p>\n\n<p>My local valid recall is 0.944 at 50 epochs.</p>",
      "rawMarkdown": "Hi all,\nIs there any way to evaluate a model at an early stage, e.g, 50 epochs? It is recommended in many posts that training 100 epochs makes the model finally converge. But it takes me 33 hours to train 100 epochs.\nIs it convincing if I compare two models' performance if they haven't finally converged?  For instance, if there're two models, one's recall is 0.945 at 50 epochs, the other's is 0.949. Can I say the second's recall is better than the first at 100 epochs?\n\nMy local valid recall is 0.944 at 50 epochs.",
      "votes": null
    },
    {
      "id": "734170",
      "postDate": "02/01/2020 03:44:30",
      "content": "<p>Probably the 0.945 model is just converging more slowly. </p>",
      "rawMarkdown": "Probably the 0.945 model is just converging more slowly.",
      "votes": null
    },
    {
      "id": "734173",
      "postDate": "02/01/2020 03:55:50",
      "content": "<p>it is like a depth first search ....</p>\n\n<ol>\n<li>train a couple of model say up to 50 epoch</li>\n<li>estimate the lower bound</li>\n<li>decide which model to continue to say 100 epoch.</li>\n<li>if results is  no good, retract to train other model</li>\n</ol>",
      "rawMarkdown": "it is like a depth first search ....\n\n1. train a couple of model say up to 50 epoch\n2. estimate the lower bound\n3. decide which model to continue to say 100 epoch.\n4. if results is  no good, retract to train other model",
      "votes": null
    },
    {
      "id": "734178",
      "postDate": "02/01/2020 03:58:18",
      "content": "<p>Thank you! Could you help me to check if there is anything wrong with my configuration?\nI use AdamW as my optimizer, <br>\nlr=0.001, \nbatch size=32, \nimage size=96x96, \nlr_scheduler is ReduceLROnPlateau (mode='min', factor=0.7, patience=5, min_lr=1e-10)\nse_resnext_32x4d as backbone with imagenet pretrained.</p>",
      "rawMarkdown": "Thank you! Could you help me to check if there is anything wrong with my configuration?\nI use AdamW as my optimizer,  \nlr=0.001, \nbatch size=32, \nimage size=96x96, \nlr_scheduler is ReduceLROnPlateau (mode='min', factor=0.7, patience=5, min_lr=1e-10)\nse_resnext_32x4d as backbone with imagenet pretrained.",
      "votes": null
    },
    {
      "id": "734197",
      "postDate": "02/01/2020 04:33:01",
      "content": "<p>😄 Thanks! You really inspire me!</p>",
      "rawMarkdown": "😄 Thanks! You really inspire me!",
      "votes": null
    },
    {
      "id": "734310",
      "postDate": "02/01/2020 08:54:56",
      "content": "<p>I know where's wrong. I add too heavy image augmentation(distortion)...😂 </p>",
      "rawMarkdown": "I know where's wrong. I add too heavy image augmentation(distortion)...😂",
      "votes": null
    },
    {
      "id": "734328",
      "postDate": "02/01/2020 09:24:47",
      "content": "<p>I get 0.943 recall score at after 3 epochs ...</p>",
      "rawMarkdown": "I get 0.943 recall score at after 3 epochs ...",
      "votes": null
    },
    {
      "id": "739328",
      "postDate": "02/07/2020 17:44:28",
      "content": "<p>this uses ray to how to do learning hyperprameter searches. It uses some smart algorithm (e.g. ASHA)to stop early if the parameters are no good</p>\n\n<p></p>\n\n<p><a href=\"https://medium.com/riselab/cutting-edge-hyperparameter-tuning-with-ray-tune-be6c0447afdf\">https://medium.com/riselab/cutting-edge-hyperparameter-tuning-with-ray-tune-be6c0447afdf</a></p>",
      "rawMarkdown": "this uses ray to how to do learning hyperprameter searches. It uses some smart algorithm (e.g. ASHA)to stop early if the parameters are no good\n\n![](https://miro.medium.com/max/640/0*ex13kSs6cKmDAIkp)\n\nhttps://medium.com/riselab/cutting-edge-hyperparameter-tuning-with-ray-tune-be6c0447afdf",
      "votes": null
    },
    {
      "id": "739823",
      "postDate": "02/08/2020 13:10:40",
      "content": "<p>another example using population based training</p>\n\n<p><a href=\"http://louiskirsch.com/ai/population-based-training\">http://louiskirsch.com/ai/population-based-training</a></p>",
      "rawMarkdown": "another example using population based training\n\nhttp://louiskirsch.com/ai/population-based-training",
      "votes": null
    },
    {
      "id": "739831",
      "postDate": "02/08/2020 13:24:26",
      "content": "<p><a href=\"/hengck23\">@hengck23</a> Thanks, Its really helpful. \nAnyway, I was wondering if you have personal blog where you keep your experimental note. 🙂 </p>",
      "rawMarkdown": "hengck23 Thanks, Its really helpful. \nAnyway, I was wondering if you have personal blog where you keep your experimental note. 🙂",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 734170,
      "author_name": "tonychenxyz",
      "author_url": "",
      "post_date": "02/01/2020 03:44:30",
      "content": "<p>Probably the 0.945 model is just converging more slowly. </p>",
      "votes": null,
      "replies": [
        {
          "id": 734178,
          "author_name": "condone",
          "author_url": "",
          "post_date": "02/01/2020 03:58:18",
          "content": "<p>Thank you! Could you help me to check if there is anything wrong with my configuration?\nI use AdamW as my optimizer, <br>\nlr=0.001, \nbatch size=32, \nimage size=96x96, \nlr_scheduler is ReduceLROnPlateau (mode='min', factor=0.7, patience=5, min_lr=1e-10)\nse_resnext_32x4d as backbone with imagenet pretrained.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 734310,
          "author_name": "condone",
          "author_url": "",
          "post_date": "02/01/2020 08:54:56",
          "content": "<p>I know where's wrong. I add too heavy image augmentation(distortion)...😂 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 734328,
          "author_name": "condone",
          "author_url": "",
          "post_date": "02/01/2020 09:24:47",
          "content": "<p>I get 0.943 recall score at after 3 epochs ...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 734173,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/01/2020 03:55:50",
      "content": "<p>it is like a depth first search ....</p>\n\n<ol>\n<li>train a couple of model say up to 50 epoch</li>\n<li>estimate the lower bound</li>\n<li>decide which model to continue to say 100 epoch.</li>\n<li>if results is  no good, retract to train other model</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 734197,
          "author_name": "condone",
          "author_url": "",
          "post_date": "02/01/2020 04:33:01",
          "content": "<p>😄 Thanks! You really inspire me!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 739328,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/07/2020 17:44:28",
      "content": "<p>this uses ray to how to do learning hyperprameter searches. It uses some smart algorithm (e.g. ASHA)to stop early if the parameters are no good</p>\n\n<p></p>\n\n<p><a href=\"https://medium.com/riselab/cutting-edge-hyperparameter-tuning-with-ray-tune-be6c0447afdf\">https://medium.com/riselab/cutting-edge-hyperparameter-tuning-with-ray-tune-be6c0447afdf</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 739823,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/08/2020 13:10:40",
          "content": "<p>another example using population based training</p>\n\n<p><a href=\"http://louiskirsch.com/ai/population-based-training\">http://louiskirsch.com/ai/population-based-training</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 739831,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "02/08/2020 13:24:26",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> Thanks, Its really helpful. \nAnyway, I was wondering if you have personal blog where you keep your experimental note. 🙂 </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "734147": "Hi all,\nIs there any way to evaluate a model at an early stage, e.g, 50 epochs? It is recommended in many posts that training 100 epochs makes the model finally converge. But it takes me 33 hours to train 100 epochs.\nIs it convincing if I compare two models' performance if they haven't finally converged?  For instance, if there're two models, one's recall is 0.945 at 50 epochs, the other's is 0.949. Can I say the second's recall is better than the first at 100 epochs?\n\nMy local valid recall is 0.944 at 50 epochs.",
    "734170": "Probably the 0.945 model is just converging more slowly.",
    "734173": "it is like a depth first search ....\n\n1. train a couple of model say up to 50 epoch\n2. estimate the lower bound\n3. decide which model to continue to say 100 epoch.\n4. if results is  no good, retract to train other model",
    "734178": "Thank you! Could you help me to check if there is anything wrong with my configuration?\nI use AdamW as my optimizer,  \nlr=0.001, \nbatch size=32, \nimage size=96x96, \nlr_scheduler is ReduceLROnPlateau (mode='min', factor=0.7, patience=5, min_lr=1e-10)\nse_resnext_32x4d as backbone with imagenet pretrained.",
    "734197": "😄 Thanks! You really inspire me!",
    "734310": "I know where's wrong. I add too heavy image augmentation(distortion)...😂",
    "734328": "I get 0.943 recall score at after 3 epochs ...",
    "739328": "this uses ray to how to do learning hyperprameter searches. It uses some smart algorithm (e.g. ASHA)to stop early if the parameters are no good\n\n![](https://miro.medium.com/max/640/0*ex13kSs6cKmDAIkp)\n\nhttps://medium.com/riselab/cutting-edge-hyperparameter-tuning-with-ray-tune-be6c0447afdf",
    "739823": "another example using population based training\n\nhttp://louiskirsch.com/ai/population-based-training",
    "739831": "hengck23 Thanks, Its really helpful. \nAnyway, I was wondering if you have personal blog where you keep your experimental note. 🙂"
  },
  "source": "meta"
}