{
  "id": 668270,
  "title": "Thank you for your participation!",
  "url": "/competitions/recodai-luc-scientific-image-forgery-detection/discussion/668270",
  "author_name": "João Phillipe Cardenuto",
  "post_date": "2026-01-15T23:58:46.073000",
  "votes": 11,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi Everyone! </p>\n<p>We are officially closing the first stage of this competition today with very positive results. Beyond the incredible scores achieved by the top solutions (superior to all methods I tested on the 1st stage test set), I am very happy to see so many participants interested in such a crucial and complex problem. </p>\n<p>We will soon start data collection for the second stage to check for the winners! </p>\n<p>For those who want to continue contributing to scientific integrity, we have launched a GitHub organization dedicated to building a community around machine learning solutions to fight fraud in science: <a href=\"https://github.com/researchintegrity\" target=\"_blank\">https://github.com/researchintegrity</a>.</p>\n<p>This is a place to centralize discussions about fraudulent cases in science, machine learning solutions to detect such cases, and other related topics. </p>\n<p>We also recently open-sourced a system dedicated to analyzing scientific images called <a href=\"https://github.com/researchintegrity/elis\" target=\"_blank\">ELIS</a>. </p>\n<p>You are all very welcome to join us there! \nOnce again, thank you for your participation! You all rock!</p>\n<p>Looking forward to the final winning solutions!</p>",
  "messages": [
    {
      "id": 3391927,
      "postDate": "2026-01-15T23:58:46.073Z",
      "content": "<p>Hi Everyone! </p>\n<p>We are officially closing the first stage of this competition today with very positive results. Beyond the incredible scores achieved by the top solutions (superior to all methods I tested on the 1st stage test set), I am very happy to see so many participants interested in such a crucial and complex problem. </p>\n<p>We will soon start data collection for the second stage to check for the winners! </p>\n<p>For those who want to continue contributing to scientific integrity, we have launched a GitHub organization dedicated to building a community around machine learning solutions to fight fraud in science: <a href=\"https://github.com/researchintegrity\" target=\"_blank\">https://github.com/researchintegrity</a>.</p>\n<p>This is a place to centralize discussions about fraudulent cases in science, machine learning solutions to detect such cases, and other related topics. </p>\n<p>We also recently open-sourced a system dedicated to analyzing scientific images called <a href=\"https://github.com/researchintegrity/elis\" target=\"_blank\">ELIS</a>. </p>\n<p>You are all very welcome to join us there! \nOnce again, thank you for your participation! You all rock!</p>\n<p>Looking forward to the final winning solutions!</p>",
      "rawMarkdown": "Hi Everyone! \n\nWe are officially closing the first stage of this competition today with very positive results. Beyond the incredible scores achieved by the top solutions (superior to all methods I tested on the 1st stage test set), I am very happy to see so many participants interested in such a crucial and complex problem. \n\nWe will soon start data collection for the second stage to check for the winners! \n\nFor those who want to continue contributing to scientific integrity, we have launched a GitHub organization dedicated to building a community around machine learning solutions to fight fraud in science: [https://github.com/researchintegrity](https://github.com/researchintegrity).\n\nThis is a place to centralize discussions about fraudulent cases in science, machine learning solutions to detect such cases, and other related topics. \n\nWe also recently open-sourced a system dedicated to analyzing scientific images called [ELIS](https://github.com/researchintegrity/elis). \n\nYou are all very welcome to join us there! \nOnce again, thank you for your participation! You all rock!\n\nLooking forward to the final winning solutions!",
      "votes": 11
    },
    {
      "id": 3392036,
      "postDate": "2026-01-16T07:21:54.520Z",
      "content": "<p>Thank you for the competition. Throughout the competition, you were both engaged and answered all the questions.</p>\n<p>I also experimented with several interesting approaches during this competition. In particular, I focused on building an agnostic pipeline that could handle both single-panel and multi-panel images within the same framework. I explored SIFT + DINO-based models in a variety of configurations and conceptual setups.</p>\n<p>One technique I applied specifically to supplementary images achieved scores around 0.55–0.60 in terms of the competition metric during my internal validation. However, when submitted to the leaderboard, the performance dropped significantly, sometimes to around 0.25. In addition to this, I conducted extensive experiments with SegFormer, training on roughly 10,000 images by incorporating multiple open-source datasets available online. Unfortunately, none of these approaches were able to match—even remotely—the scores achieved by the public notebooks.</p>\n<p>Based on my observations, I suspect that solutions scoring 0.36 and above on the LB are able to handle multi-panel images to some extent. Despite substantial effort, I wasn’t able to reach the performance level I had hoped for. My hypothesis is that high-scoring solutions may be particularly effective on certain types of forgeries; if the newly introduced images differ from those high-frequency forgery patterns, their performance might degrade. Overall, I feel there may be some mismatch between the data distribution and the intended expectations of the task.</p>\n<p>I am very curious to see the winning solutions and sincerely hope they will be shared once the competition concludes.</p>\n<p>Finally, the ELIS system looks extremely impressive and clearly reflects many years of dedicated R&amp;D. Congratulations to the researchers and engineers behind it.</p>",
      "rawMarkdown": "Thank you for the competition. Throughout the competition, you were both engaged and answered all the questions.\n\nI also experimented with several interesting approaches during this competition. In particular, I focused on building an agnostic pipeline that could handle both single-panel and multi-panel images within the same framework. I explored SIFT + DINO-based models in a variety of configurations and conceptual setups.\n\nOne technique I applied specifically to supplementary images achieved scores around 0.55–0.60 in terms of the competition metric during my internal validation. However, when submitted to the leaderboard, the performance dropped significantly, sometimes to around 0.25. In addition to this, I conducted extensive experiments with SegFormer, training on roughly 10,000 images by incorporating multiple open-source datasets available online. Unfortunately, none of these approaches were able to match—even remotely—the scores achieved by the public notebooks.\n\nBased on my observations, I suspect that solutions scoring 0.36 and above on the LB are able to handle multi-panel images to some extent. Despite substantial effort, I wasn’t able to reach the performance level I had hoped for. My hypothesis is that high-scoring solutions may be particularly effective on certain types of forgeries; if the newly introduced images differ from those high-frequency forgery patterns, their performance might degrade. Overall, I feel there may be some mismatch between the data distribution and the intended expectations of the task.\n\nI am very curious to see the winning solutions and sincerely hope they will be shared once the competition concludes.\n\nFinally, the ELIS system looks extremely impressive and clearly reflects many years of dedicated R&D. Congratulations to the researchers and engineers behind it.",
      "votes": 1,
      "replies": [
        {
          "id": 3392229,
          "postDate": "2026-01-16T14:00:35.363Z",
          "content": "<p>Yes, instead of sift+dino, if you did sift + detector extractor (like Yolo or sam) then that started to boozt the score.</p>",
          "rawMarkdown": "Yes, instead of sift+dino, if you did sift + detector extractor (like Yolo or sam) then that started to boozt the score.",
          "votes": 1,
          "replies": [
            {
              "id": 3392246,
              "postDate": "2026-01-16T14:49:00.673Z",
              "content": "<p>You're right. \nActually, I tried some really good techniques. I separated the images into single/multi panels. I worked specifically on the multi panels. I separated them into panels using YOLO and worked on the panels individually. The test results were good too. But I couldn't get the score I wanted on LB.</p>",
              "rawMarkdown": "You're right. \nActually, I tried some really good techniques. I separated the images into single/multi panels. I worked specifically on the multi panels. I separated them into panels using YOLO and worked on the panels individually. The test results were good too. But I couldn't get the score I wanted on LB.",
              "votes": 1
            }
          ]
        },
        {
          "id": 3392794,
          "postDate": "2026-01-17T14:18:08.270Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/musapeker\" target=\"_blank\">@musapeker</a> ,</p>\n<p>Thank you for sharing your experience with the competition and for your kind words!</p>\n<p>When designing the competition, I also believed that a multi-panel parser combined with a clone detector would be key to achieving better results. I noticed multiple participants trying to use image segmentation approaches for copy-move detection. In my own experience, I also struggled to find a robust solution using that method. The models usually overfit to specific image types, which is a major challenge given that the biomedical domain is so diverse.</p>\n<p>I am really glad to hear you checked out ELIS!</p>\n<p>Thanks again for your participation and good luck in stage 2!</p>",
          "rawMarkdown": "Hi @musapeker ,\n\nThank you for sharing your experience with the competition and for your kind words!\n\nWhen designing the competition, I also believed that a multi-panel parser combined with a clone detector would be key to achieving better results. I noticed multiple participants trying to use image segmentation approaches for copy-move detection. In my own experience, I also struggled to find a robust solution using that method. The models usually overfit to specific image types, which is a major challenge given that the biomedical domain is so diverse.\n\nI am really glad to hear you checked out ELIS!\n\nThanks again for your participation and good luck in stage 2!",
          "votes": 1
        }
      ]
    },
    {
      "id": 3399193,
      "postDate": "2026-01-30T11:28:11.940Z",
      "content": "<p>Hi, great competition.\nWhy is it still showing \"3 months to go\" with submission disabled?</p>",
      "rawMarkdown": "Hi, great competition.\nWhy is it still showing \"3 months to go\" with submission disabled?"
    },
    {
      "id": 3391987,
      "postDate": "2026-01-16T03:59:41.223Z",
      "content": "<p>Thank you for the competition. I spent a lot of time on it and luckily did well in the Public LB; hopefully it translates to the Stage 2 LB.</p>\n<p>1) Throughout this competition, I saw many software vendors selling Copy-Forge detection. I thought it seemed a solved problem, you just had to pay for it. And it seemed this competition was basically to have Kagglers create a free version that matched their quality (beyond what Forensically can show you). Was that part of the goal of this competition?\n2) Why focus just on copy-forge detection, instead of also on manipulation, or even AI generated image? Is copy-forge the task that ELIS struggles most on?\n3) I created a handlabeled dataset of around 500 images from RetractionWatch containing copy-forge duplicates if that’s of interest. Unfortunately, even training NN on it, I was not able to get much use out of it. The numpy masks are in the same format as the training set. Would you like me to share it?\n4) What is the ultimate desire - to have all journals adopt ELIS before accepting a paper, so they knoe there are no mistakes?</p>",
      "rawMarkdown": "Thank you for the competition. I spent a lot of time on it and luckily did well in the Public LB; hopefully it translates to the Stage 2 LB.\n\n1) Throughout this competition, I saw many software vendors selling Copy-Forge detection. I thought it seemed a solved problem, you just had to pay for it. And it seemed this competition was basically to have Kagglers create a free version that matched their quality (beyond what Forensically can show you). Was that part of the goal of this competition?\n2) Why focus just on copy-forge detection, instead of also on manipulation, or even AI generated image? Is copy-forge the task that ELIS struggles most on?\n3) I created a handlabeled dataset of around 500 images from RetractionWatch containing copy-forge duplicates if that’s of interest. Unfortunately, even training NN on it, I was not able to get much use out of it. The numpy masks are in the same format as the training set. Would you like me to share it?\n4) What is the ultimate desire - to have all journals adopt ELIS before accepting a paper, so they knoe there are no mistakes?",
      "replies": [
        {
          "id": 3391988,
          "postDate": "2026-01-16T04:08:13.367Z",
          "content": "<p>My solution is quite similar to ELIS (extraction-&gt;keypoint matching). I also tried approach where instead of pure keypoint, have NN be fed two panels and output the segmentation where they were duplicated, but I did not have time to submit this. In general I found NN performance could not beat traditional keypoint methods. It seems Public Notebook DINO models just highlighting Western Blots which happen to have higher frequency of forgery, so they do better.</p>",
          "rawMarkdown": "My solution is quite similar to ELIS (extraction->keypoint matching). I also tried approach where instead of pure keypoint, have NN be fed two panels and output the segmentation where they were duplicated, but I did not have time to submit this. In general I found NN performance could not beat traditional keypoint methods. It seems Public Notebook DINO models just highlighting Western Blots which happen to have higher frequency of forgery, so they do better.",
          "votes": 2
        },
        {
          "id": 3392800,
          "postDate": "2026-01-17T14:25:26.357Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/returnofsputnik\" target=\"_blank\">@returnofsputnik</a> ,</p>\n<p>Thanks for sharing your thoughts and for the time you dedicated to the competition!</p>\n<p>1- When I first started working on this problem back in 2018, I had the same impression as you, it seemed like a \"solved problem.\"</p>\n<p>However, when we tested the best available detectors on biomedical images (especially those extracted from scientific papers) none of them worked reliably. One of the biggest blockers has been the lack of a proper benchmark with real cases. Gathering this data is difficult due to paywalls and copyright nuances, which we had to navigate carefully to make this competition possible. While I am aware of commercial solutions, they lack a standard benchmark, so their performance on this specific domain is unclear to our knowledge. Also, unfortunately, these tools are often inaccessible to many researchers and institutions (including us here in Latin America). So, a major goal of this competition is to facilitate research and democratize access to a transparent benchmark and tool for the problem.</p>\n<p>2- We focused on copy-move because it remains a recurrently reported issue in science and technically remains an open problem. While we are certainly concerned about AI-generated images, there are currently very few reported cases of them (which may be due to underreporting or difficulty in detection). It would have been nearly impossible to build a robust dataset based on real cases for AI-generated images right now, but that could be a nice future competition to plan :)</p>\n<p>3- Regarding your dataset, any new dataset is very welcome! If you decide to share it, please just be careful to check the image licenses and any legal restrictions before making it public on Kaggle (you can also contact me directly to share it).</p>\n<p>4- Today, the detection of fraudulent cases relies heavily on the courage of independent and free work of whistleblowers to spot manipulated images, most of which have already passed peer review. The goal of ELIS is to democratize forensics tools so they can be used by anyone, whether that is a publisher, a university integrity office, or an independent whistleblower. </p>\n<p>Our goal with ELIS is to shed light on the problem and foster more research and solutions. My experience with the problem says that it won't be solved by a single solution or entity, but by a diverse, multi-disciplinary community working continuously. This is especially critical now that we are seeing organizations, known as \"paper mills\",  dedicated to producing fraudulent papers ( <a href=\"https://www.nature.com/articles/d41586-025-01824-3\" target=\"_blank\">https://www.nature.com/articles/d41586-025-01824-3</a>, <a href=\"https://www.nature.com/articles/d41586-021-00733-5\" target=\"_blank\">https://www.nature.com/articles/d41586-021-00733-5</a>).</p>\n<p>Good luck in Stage 2!</p>",
          "rawMarkdown": "Hi @returnofsputnik ,\n\nThanks for sharing your thoughts and for the time you dedicated to the competition!\n\n1- When I first started working on this problem back in 2018, I had the same impression as you, it seemed like a \"solved problem.\"\n\nHowever, when we tested the best available detectors on biomedical images (especially those extracted from scientific papers) none of them worked reliably. One of the biggest blockers has been the lack of a proper benchmark with real cases. Gathering this data is difficult due to paywalls and copyright nuances, which we had to navigate carefully to make this competition possible. While I am aware of commercial solutions, they lack a standard benchmark, so their performance on this specific domain is unclear to our knowledge. Also, unfortunately, these tools are often inaccessible to many researchers and institutions (including us here in Latin America). So, a major goal of this competition is to facilitate research and democratize access to a transparent benchmark and tool for the problem.\n\n2- We focused on copy-move because it remains a recurrently reported issue in science and technically remains an open problem. While we are certainly concerned about AI-generated images, there are currently very few reported cases of them (which may be due to underreporting or difficulty in detection). It would have been nearly impossible to build a robust dataset based on real cases for AI-generated images right now, but that could be a nice future competition to plan :)\n\n3- Regarding your dataset, any new dataset is very welcome! If you decide to share it, please just be careful to check the image licenses and any legal restrictions before making it public on Kaggle (you can also contact me directly to share it).\n\n4- Today, the detection of fraudulent cases relies heavily on the courage of independent and free work of whistleblowers to spot manipulated images, most of which have already passed peer review. The goal of ELIS is to democratize forensics tools so they can be used by anyone, whether that is a publisher, a university integrity office, or an independent whistleblower. \n\nOur goal with ELIS is to shed light on the problem and foster more research and solutions. My experience with the problem says that it won't be solved by a single solution or entity, but by a diverse, multi-disciplinary community working continuously. This is especially critical now that we are seeing organizations, known as \"paper mills\",  dedicated to producing fraudulent papers ( [https://www.nature.com/articles/d41586-025-01824-3](https://www.nature.com/articles/d41586-025-01824-3), [https://www.nature.com/articles/d41586-021-00733-5](https://www.nature.com/articles/d41586-021-00733-5)).\n\nGood luck in Stage 2!"
        }
      ]
    },
    {
      "id": 3395529,
      "postDate": "2026-01-23T06:21:53.067Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3392585,
      "postDate": "2026-01-17T07:08:54.637Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 3392792,
          "postDate": "2026-01-17T14:16:10.753Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/kami1976\" target=\"_blank\">@kami1976</a> ,</p>\n<p>After the 1st stage of the competition closed, Kaggle Staff conducted a submission review and removed participants who were found to be in violation of the rules from the leaderboard.</p>\n<p>This removal is often due to the detection of multiple accounts (used to bypass daily submission limits) or other rule violations.</p>\n<p>As competition hosts, we do not control these removals. If you believe a mistake has been made, I recommend checking the competition rules and contacting Kaggle Support directly to appeal.</p>",
          "rawMarkdown": "Hi @kami1976 ,\n\nAfter the 1st stage of the competition closed, Kaggle Staff conducted a submission review and removed participants who were found to be in violation of the rules from the leaderboard.\n\nThis removal is often due to the detection of multiple accounts (used to bypass daily submission limits) or other rule violations.\n\nAs competition hosts, we do not control these removals. If you believe a mistake has been made, I recommend checking the competition rules and contacting Kaggle Support directly to appeal.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3392036,
      "author_name": "Musa Peker",
      "author_url": "",
      "post_date": "2026-01-16T07:21:54.520000",
      "content": "<p>Thank you for the competition. Throughout the competition, you were both engaged and answered all the questions.</p>\n<p>I also experimented with several interesting approaches during this competition. In particular, I focused on building an agnostic pipeline that could handle both single-panel and multi-panel images within the same framework. I explored SIFT + DINO-based models in a variety of configurations and conceptual setups.</p>\n<p>One technique I applied specifically to supplementary images achieved scores around 0.55–0.60 in terms of the competition metric during my internal validation. However, when submitted to the leaderboard, the performance dropped significantly, sometimes to around 0.25. In addition to this, I conducted extensive experiments with SegFormer, training on roughly 10,000 images by incorporating multiple open-source datasets available online. Unfortunately, none of these approaches were able to match—even remotely—the scores achieved by the public notebooks.</p>\n<p>Based on my observations, I suspect that solutions scoring 0.36 and above on the LB are able to handle multi-panel images to some extent. Despite substantial effort, I wasn’t able to reach the performance level I had hoped for. My hypothesis is that high-scoring solutions may be particularly effective on certain types of forgeries; if the newly introduced images differ from those high-frequency forgery patterns, their performance might degrade. Overall, I feel there may be some mismatch between the data distribution and the intended expectations of the task.</p>\n<p>I am very curious to see the winning solutions and sincerely hope they will be shared once the competition concludes.</p>\n<p>Finally, the ELIS system looks extremely impressive and clearly reflects many years of dedicated R&amp;D. Congratulations to the researchers and engineers behind it.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3392229,
          "author_name": "CoreyJamesLevinson",
          "author_url": "",
          "post_date": "2026-01-16T14:00:35.363000",
          "content": "<p>Yes, instead of sift+dino, if you did sift + detector extractor (like Yolo or sam) then that started to boozt the score.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3392246,
              "author_name": "Musa Peker",
              "author_url": "",
              "post_date": "2026-01-16T14:49:00.673000",
              "content": "<p>You're right. \nActually, I tried some really good techniques. I separated the images into single/multi panels. I worked specifically on the multi panels. I separated them into panels using YOLO and worked on the panels individually. The test results were good too. But I couldn't get the score I wanted on LB.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 3392794,
          "author_name": "João Phillipe Cardenuto",
          "author_url": "",
          "post_date": "2026-01-17T14:18:08.270000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/musapeker\" target=\"_blank\">@musapeker</a> ,</p>\n<p>Thank you for sharing your experience with the competition and for your kind words!</p>\n<p>When designing the competition, I also believed that a multi-panel parser combined with a clone detector would be key to achieving better results. I noticed multiple participants trying to use image segmentation approaches for copy-move detection. In my own experience, I also struggled to find a robust solution using that method. The models usually overfit to specific image types, which is a major challenge given that the biomedical domain is so diverse.</p>\n<p>I am really glad to hear you checked out ELIS!</p>\n<p>Thanks again for your participation and good luck in stage 2!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3399193,
      "author_name": "Tashi Jawed",
      "author_url": "",
      "post_date": "2026-01-30T11:28:11.940000",
      "content": "<p>Hi, great competition.\nWhy is it still showing \"3 months to go\" with submission disabled?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3391987,
      "author_name": "CoreyJamesLevinson",
      "author_url": "",
      "post_date": "2026-01-16T03:59:41.223000",
      "content": "<p>Thank you for the competition. I spent a lot of time on it and luckily did well in the Public LB; hopefully it translates to the Stage 2 LB.</p>\n<p>1) Throughout this competition, I saw many software vendors selling Copy-Forge detection. I thought it seemed a solved problem, you just had to pay for it. And it seemed this competition was basically to have Kagglers create a free version that matched their quality (beyond what Forensically can show you). Was that part of the goal of this competition?\n2) Why focus just on copy-forge detection, instead of also on manipulation, or even AI generated image? Is copy-forge the task that ELIS struggles most on?\n3) I created a handlabeled dataset of around 500 images from RetractionWatch containing copy-forge duplicates if that’s of interest. Unfortunately, even training NN on it, I was not able to get much use out of it. The numpy masks are in the same format as the training set. Would you like me to share it?\n4) What is the ultimate desire - to have all journals adopt ELIS before accepting a paper, so they knoe there are no mistakes?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3391988,
          "author_name": "CoreyJamesLevinson",
          "author_url": "",
          "post_date": "2026-01-16T04:08:13.367000",
          "content": "<p>My solution is quite similar to ELIS (extraction-&gt;keypoint matching). I also tried approach where instead of pure keypoint, have NN be fed two panels and output the segmentation where they were duplicated, but I did not have time to submit this. In general I found NN performance could not beat traditional keypoint methods. It seems Public Notebook DINO models just highlighting Western Blots which happen to have higher frequency of forgery, so they do better.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 3392800,
          "author_name": "João Phillipe Cardenuto",
          "author_url": "",
          "post_date": "2026-01-17T14:25:26.357000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/returnofsputnik\" target=\"_blank\">@returnofsputnik</a> ,</p>\n<p>Thanks for sharing your thoughts and for the time you dedicated to the competition!</p>\n<p>1- When I first started working on this problem back in 2018, I had the same impression as you, it seemed like a \"solved problem.\"</p>\n<p>However, when we tested the best available detectors on biomedical images (especially those extracted from scientific papers) none of them worked reliably. One of the biggest blockers has been the lack of a proper benchmark with real cases. Gathering this data is difficult due to paywalls and copyright nuances, which we had to navigate carefully to make this competition possible. While I am aware of commercial solutions, they lack a standard benchmark, so their performance on this specific domain is unclear to our knowledge. Also, unfortunately, these tools are often inaccessible to many researchers and institutions (including us here in Latin America). So, a major goal of this competition is to facilitate research and democratize access to a transparent benchmark and tool for the problem.</p>\n<p>2- We focused on copy-move because it remains a recurrently reported issue in science and technically remains an open problem. While we are certainly concerned about AI-generated images, there are currently very few reported cases of them (which may be due to underreporting or difficulty in detection). It would have been nearly impossible to build a robust dataset based on real cases for AI-generated images right now, but that could be a nice future competition to plan :)</p>\n<p>3- Regarding your dataset, any new dataset is very welcome! If you decide to share it, please just be careful to check the image licenses and any legal restrictions before making it public on Kaggle (you can also contact me directly to share it).</p>\n<p>4- Today, the detection of fraudulent cases relies heavily on the courage of independent and free work of whistleblowers to spot manipulated images, most of which have already passed peer review. The goal of ELIS is to democratize forensics tools so they can be used by anyone, whether that is a publisher, a university integrity office, or an independent whistleblower. </p>\n<p>Our goal with ELIS is to shed light on the problem and foster more research and solutions. My experience with the problem says that it won't be solved by a single solution or entity, but by a diverse, multi-disciplinary community working continuously. This is especially critical now that we are seeing organizations, known as \"paper mills\",  dedicated to producing fraudulent papers ( <a href=\"https://www.nature.com/articles/d41586-025-01824-3\" target=\"_blank\">https://www.nature.com/articles/d41586-025-01824-3</a>, <a href=\"https://www.nature.com/articles/d41586-021-00733-5\" target=\"_blank\">https://www.nature.com/articles/d41586-021-00733-5</a>).</p>\n<p>Good luck in Stage 2!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3395529,
      "author_name": "",
      "author_url": "",
      "post_date": "2026-01-23T06:21:53.067000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3392585,
      "author_name": "",
      "author_url": "",
      "post_date": "2026-01-17T07:08:54.637000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 3392792,
          "author_name": "João Phillipe Cardenuto",
          "author_url": "",
          "post_date": "2026-01-17T14:16:10.753000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/kami1976\" target=\"_blank\">@kami1976</a> ,</p>\n<p>After the 1st stage of the competition closed, Kaggle Staff conducted a submission review and removed participants who were found to be in violation of the rules from the leaderboard.</p>\n<p>This removal is often due to the detection of multiple accounts (used to bypass daily submission limits) or other rule violations.</p>\n<p>As competition hosts, we do not control these removals. If you believe a mistake has been made, I recommend checking the competition rules and contacting Kaggle Support directly to appeal.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3391927": "Hi Everyone! \n\nWe are officially closing the first stage of this competition today with very positive results. Beyond the incredible scores achieved by the top solutions (superior to all methods I tested on the 1st stage test set), I am very happy to see so many participants interested in such a crucial and complex problem. \n\nWe will soon start data collection for the second stage to check for the winners! \n\nFor those who want to continue contributing to scientific integrity, we have launched a GitHub organization dedicated to building a community around machine learning solutions to fight fraud in science: [https://github.com/researchintegrity](https://github.com/researchintegrity).\n\nThis is a place to centralize discussions about fraudulent cases in science, machine learning solutions to detect such cases, and other related topics. \n\nWe also recently open-sourced a system dedicated to analyzing scientific images called [ELIS](https://github.com/researchintegrity/elis). \n\nYou are all very welcome to join us there! \nOnce again, thank you for your participation! You all rock!\n\nLooking forward to the final winning solutions!",
    "3392036": "Thank you for the competition. Throughout the competition, you were both engaged and answered all the questions.\n\nI also experimented with several interesting approaches during this competition. In particular, I focused on building an agnostic pipeline that could handle both single-panel and multi-panel images within the same framework. I explored SIFT + DINO-based models in a variety of configurations and conceptual setups.\n\nOne technique I applied specifically to supplementary images achieved scores around 0.55–0.60 in terms of the competition metric during my internal validation. However, when submitted to the leaderboard, the performance dropped significantly, sometimes to around 0.25. In addition to this, I conducted extensive experiments with SegFormer, training on roughly 10,000 images by incorporating multiple open-source datasets available online. Unfortunately, none of these approaches were able to match—even remotely—the scores achieved by the public notebooks.\n\nBased on my observations, I suspect that solutions scoring 0.36 and above on the LB are able to handle multi-panel images to some extent. Despite substantial effort, I wasn’t able to reach the performance level I had hoped for. My hypothesis is that high-scoring solutions may be particularly effective on certain types of forgeries; if the newly introduced images differ from those high-frequency forgery patterns, their performance might degrade. Overall, I feel there may be some mismatch between the data distribution and the intended expectations of the task.\n\nI am very curious to see the winning solutions and sincerely hope they will be shared once the competition concludes.\n\nFinally, the ELIS system looks extremely impressive and clearly reflects many years of dedicated R&D. Congratulations to the researchers and engineers behind it.",
    "3399193": "Hi, great competition.\nWhy is it still showing \"3 months to go\" with submission disabled?",
    "3391987": "Thank you for the competition. I spent a lot of time on it and luckily did well in the Public LB; hopefully it translates to the Stage 2 LB.\n\n1) Throughout this competition, I saw many software vendors selling Copy-Forge detection. I thought it seemed a solved problem, you just had to pay for it. And it seemed this competition was basically to have Kagglers create a free version that matched their quality (beyond what Forensically can show you). Was that part of the goal of this competition?\n2) Why focus just on copy-forge detection, instead of also on manipulation, or even AI generated image? Is copy-forge the task that ELIS struggles most on?\n3) I created a handlabeled dataset of around 500 images from RetractionWatch containing copy-forge duplicates if that’s of interest. Unfortunately, even training NN on it, I was not able to get much use out of it. The numpy masks are in the same format as the training set. Would you like me to share it?\n4) What is the ultimate desire - to have all journals adopt ELIS before accepting a paper, so they knoe there are no mistakes?",
    "3395529": "",
    "3392585": ""
  }
}