{
  "id": 346372,
  "title": "Question about SEED fixing",
  "url": "/competitions/amex-default-prediction/discussion/346372",
  "author_name": "",
  "post_date": "2022-08-19T06:36:46.996307300Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I encountered a curious bug when I started developing NN models using Pytorch.<br>\nFirst, in the first line below, I set random seeds as usual so that I can get reproductive results followed by nomal MLP model definition, training, and validation. <br>\nHowever, I realized the result decreases to some extent when I don't set the <code>SECOND</code> \"seed_everything(config.SEED)\" function and I ended up not being able to reproduce the same cv score during training and validation.<br>\nI checked that I set model.eval() properly.<br>\nDoes anyone know why I have to set random seeds twice in this case?<br>\nThanks in advance !</p>\n<pre><code>seed_everything(config.SEED) \n-------- MODEL &amp; CV DEFINITION -----------\n-------- TRAINING &amp; VALIDATION ---------\nBelow are inference codes to get the best cv score during the training\n\ndef predict(model, dataloader, device='cpu'):\n    model.eval()\n    outputs = []\n    for item in tqdm(dataloader, total=len(dataloader)):\n        num_feats = item[0].to(device).float()\n        output = model(num_feats)\n        outputs.extend(output.view(-1).data.cpu().numpy())\n    return outputs\n\nseed_everything(config.SEED)  #Second seed fixing. This is somehow required to get the same results.\nbest_model = AmexModel()\nbest_model.load_state_dict(torch.load(config.save_path + f\"/amex_model_fold{fold}.pt\"))\nbest_model.to(device)\noof_preds = predict(best_model, valid_dataloader, device)\nacc = amex_metric(Xy_valid['target'].values, oof_preds)\nprint('Kaggle Metric =',acc,'\\n')\n</code></pre>",
  "messages": [
    {
      "id": "1905580",
      "postDate": "08/19/2022 06:36:46",
      "content": "<p>I encountered a curious bug when I started developing NN models using Pytorch.<br>\nFirst, in the first line below, I set random seeds as usual so that I can get reproductive results followed by nomal MLP model definition, training, and validation. <br>\nHowever, I realized the result decreases to some extent when I don't set the <code>SECOND</code> \"seed_everything(config.SEED)\" function and I ended up not being able to reproduce the same cv score during training and validation.<br>\nI checked that I set model.eval() properly.<br>\nDoes anyone know why I have to set random seeds twice in this case?<br>\nThanks in advance !</p>\n<pre><code>seed_everything(config.SEED) \n-------- MODEL &amp; CV DEFINITION -----------\n-------- TRAINING &amp; VALIDATION ---------\nBelow are inference codes to get the best cv score during the training\n\ndef predict(model, dataloader, device='cpu'):\n    model.eval()\n    outputs = []\n    for item in tqdm(dataloader, total=len(dataloader)):\n        num_feats = item[0].to(device).float()\n        output = model(num_feats)\n        outputs.extend(output.view(-1).data.cpu().numpy())\n    return outputs\n\nseed_everything(config.SEED)  #Second seed fixing. This is somehow required to get the same results.\nbest_model = AmexModel()\nbest_model.load_state_dict(torch.load(config.save_path + f\"/amex_model_fold{fold}.pt\"))\nbest_model.to(device)\noof_preds = predict(best_model, valid_dataloader, device)\nacc = amex_metric(Xy_valid['target'].values, oof_preds)\nprint('Kaggle Metric =',acc,'\\n')\n</code></pre>",
      "rawMarkdown": "I encountered a curious bug when I started developing NN models using Pytorch.\nFirst, in the first line below, I set random seeds as usual so that I can get reproductive results followed by nomal MLP model definition, training, and validation. \nHowever, I realized the result decreases to some extent when I don't set the `SECOND` \"seed_everything(config.SEED)\" function and I ended up not being able to reproduce the same cv score during training and validation.\nI checked that I set model.eval() properly.\nDoes anyone know why I have to set random seeds twice in this case?\nThanks in advance !\n\n```\nseed_everything(config.SEED) \n-------- MODEL & CV DEFINITION -----------\n-------- TRAINING & VALIDATION ---------\nBelow are inference codes to get the best cv score during the training\n\ndef predict(model, dataloader, device='cpu'):\n    model.eval()\n    outputs = []\n    for item in tqdm(dataloader, total=len(dataloader)):\n        num_feats = item[0].to(device).float()\n        output = model(num_feats)\n        outputs.extend(output.view(-1).data.cpu().numpy())\n    return outputs\n\nseed_everything(config.SEED)  #Second seed fixing. This is somehow required to get the same results.\nbest_model = AmexModel()\nbest_model.load_state_dict(torch.load(config.save_path + f\"/amex_model_fold{fold}.pt\"))\nbest_model.to(device)\noof_preds = predict(best_model, valid_dataloader, device)\nacc = amex_metric(Xy_valid['target'].values, oof_preds)\nprint('Kaggle Metric =',acc,'\\n')\n```",
      "votes": null
    },
    {
      "id": "1905718",
      "postDate": "08/19/2022 08:46:42",
      "content": "<p>Maybe you are importing or initializing something inbetween the 1st seed_everything and 2nd which is setting its own seed, thus adding the 2nd is needing to set the seeds on those imported/initialized items..?</p>",
      "rawMarkdown": "Maybe you are importing or initializing something inbetween the 1st seed_everything and 2nd which is setting its own seed, thus adding the 2nd is needing to set the seeds on those imported/initialized items..?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1905718,
      "author_name": "julianmukaj",
      "author_url": "",
      "post_date": "08/19/2022 08:46:42",
      "content": "<p>Maybe you are importing or initializing something inbetween the 1st seed_everything and 2nd which is setting its own seed, thus adding the 2nd is needing to set the seeds on those imported/initialized items..?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1905580": "I encountered a curious bug when I started developing NN models using Pytorch.\nFirst, in the first line below, I set random seeds as usual so that I can get reproductive results followed by nomal MLP model definition, training, and validation. \nHowever, I realized the result decreases to some extent when I don't set the `SECOND` \"seed_everything(config.SEED)\" function and I ended up not being able to reproduce the same cv score during training and validation.\nI checked that I set model.eval() properly.\nDoes anyone know why I have to set random seeds twice in this case?\nThanks in advance !\n\n```\nseed_everything(config.SEED) \n-------- MODEL & CV DEFINITION -----------\n-------- TRAINING & VALIDATION ---------\nBelow are inference codes to get the best cv score during the training\n\ndef predict(model, dataloader, device='cpu'):\n    model.eval()\n    outputs = []\n    for item in tqdm(dataloader, total=len(dataloader)):\n        num_feats = item[0].to(device).float()\n        output = model(num_feats)\n        outputs.extend(output.view(-1).data.cpu().numpy())\n    return outputs\n\nseed_everything(config.SEED)  #Second seed fixing. This is somehow required to get the same results.\nbest_model = AmexModel()\nbest_model.load_state_dict(torch.load(config.save_path + f\"/amex_model_fold{fold}.pt\"))\nbest_model.to(device)\noof_preds = predict(best_model, valid_dataloader, device)\nacc = amex_metric(Xy_valid['target'].values, oof_preds)\nprint('Kaggle Metric =',acc,'\\n')\n```",
    "1905718": "Maybe you are importing or initializing something inbetween the 1st seed_everything and 2nd which is setting its own seed, thus adding the 2nd is needing to set the seeds on those imported/initialized items..?"
  },
  "source": "meta"
}