Biomedical image analysis competitions: The state of current participation practice

M. Eisenmann, A. Reinke, V. Weru, M. Tizabi, F. Isensee, T. Adler, P. Godau, V. Cheplygina, M. Kozubek, S. Ali, A. Gupta, J. Kybic, A. Noble, C. de Solórzano, S. Pachade, C. Petitjean, D. Sage, D. Wei, E. Wilden, D. Alapatt, V. Andrearczyk, U. Baid, S. Bakas, N. Balu, S. Bano, V. Bawa, J. Bernal, S. Bodenstedt, A. Casella, J. Choi, O. Commowick, M. Daum, A. Depeursinge, R. Dorent, J. Egger, H. Eichhorn, S. Engelhardt, M. Ganz, G. Girard, L. Hansen, M. Heinrich, N. Heller, A. Hering, A. Huaulmé, H. Kim, B. Landman, H. Li, J. Li, J. Ma, A. Martel, C. Martín-Isla, B. Menze, C. Nwoye, V. Oreiller, N. Padoy, S. Pati, K. Payette, C. Sudre, K. van Wijnen, A. Vardazaryan, T. Vercauteren, M. Wagner, C. Wang, M. Yap, Z. Yu, C. Yuan, M. Zenk, A. Zia, D. Zimmerer, R. Bao, C. Choi, A. Cohen, O. Dzyubachyk, A. Galdran, T. Gan, T. Guo, P. Gupta, M. Haithami, E. Ho, I. Jang, Z. Li, Z. Luo, F. Lux, S. Makrogiannis, D. Müller, Y. Oh, S. Pang, C. Pape, G. Polat, C. Reed, K. Ryu, T. Scherr, V. Thambawita, H. Wang, X. Wang, K. Xu, H. Yeh, D. Yeo, Y. Yuan, Y. Zeng, X. Zhao, J. Abbing, J. Adam, N. Adluru, N. Agethen, S. Ahmed, Y. Khalil, M. Alenyà, E. Alhoniemi, C. An, T. Anwar, T. Arega, N. Avisdris, D. Aydogan, Y. Bai, M. Calisto, B. Basaran, M. Beetz, C. Bian, H. Bian, K. Blansit, L. Bloch, R. Bohnsack, S. Bosticardo, J. Breen, M. Brudfors, R. Brüngel, M. Cabezas, A. Cacciola, Z. Chen, Y. Chen, D. Chen, M. Cho, M. Choi, C. Xie, D. Cobzas, J. Cohen-Adad, J. Acero, S. Das, M. de Oliveira, H. Deng, G. Dong, L. Doorenbos, C. Efird, S. Escalera, D. Fan, M. Serj, A. Fenneteau, L. Fidon, P. Filipiak, R. Finzel, N. Freitas, C. Friedrich, M. Fulton, F. Gaida, F. Galati, C. Galazis, C. Gan, Z. Gao, S. Gao, M. Gazda, B. Gerats, N. Getty, A. Gibicar, R. Gifford, S. Gohil, M. Grammatikopoulou, D. Grzech, O. Güley, T. Günnemann, C. Guo, S. Guy, H. Ha, L. Han, I. Han, A. Hatamizadeh, T. He, J. Heo, S. Hitziger, S. Hong, R. Huang, Z. Huang, M. Huellebrand, S. Huschauer, M. Hussain, T. Inubushi, E. Polat, M. Jafaritadi, S. Jeong, B. Jian, Y. Jiang, Z. Jiang, Y. Jin, S. Joshi, A. Kadkhodamohammadi, R. Kamraoui, I. Kang, J. Kang, D. Karimi, A. Khademi, M. Khan, S. Khan, R. Khantwal, K. Kim, T. Kline, S. Kondo, E. Kontio, A. Krenzer, A. Kroviakov, H. Kuijf, S. Kumar, F. La Rosa, A. Lad, D. Lee, M. Lee, C. Lena, L. Li, X. Li, F. Liao, K. Liao, A. Oliveira, C. Lin, S. Lin, A. Linardos, M. Linguraru, H. Liu, T. Liu, D. Liu, Y. Liu, J. Lourenço-Silva, J. Lu, I. Luengo, C. Lund, H. Luu, Y. Lv, U. Macar, L. Maechler, S. L., K. Marshall, M. Mazher, R. McKinley, A. Medela, F. Meissen, M. Meng, D. Miller, S. Mirjahanmardi, A. Mishra, S. Mitha, H. Mohy-ud-Din, T. Mok, G. Murugesan, E. Karthik, S. Nalawade, J. Nalepa, M. Naser, R. Nateghi, H. Naveed, Q. Nguyen, C. Quoc, B. Nichyporuk, B. Oliveira, D. Owen, J. Pal, J. Pan, W. Pan, W. Pang, B. Park, V. Pawar, K. Pawar, M. Peven, L. Philipp, T. Pieciak, S. Plotka, M. Plutat, F. Pourakpour, D. Preloznik, K. Punithakumar, A. Qayyum, S. Queirós, A. Rahmim, S. Razavi, J. Ren, M. Rezaei, J. Rico, Z. Rieu, M. Rink, J. Roth, Y. Ruiz-Gonzalez, N. Saeed, A. Saha, M. Salem, R. Sanchez-Matilla, K. Schilling, W. Shao, Z. Shen, R. Shi, P. Shi, D. Sobotka, T. Soulier, B. Fadida, D. Stoyanov, T. Mun, X. Sun, R. Tao, F. Thaler, A. Théberge, F. Thielke, H. Torres, K. Wahid, J. Wang, Y. Wang, W. Wang, J. Wen, N. Wen, M. Wodzinski, Y. Wu, F. Xia, T. Xiang, C. Xiaofei, L. Xu, T. Xue, Y. Yang, L. Yang, K. Yao, H. Yao, A. Yazdani, M. Yip, H. Yoo, F. Yousefirizi, S. Yu, L. Yu, J. Zamora, R. Zeineldin, D. Zeng, J. Zhang, B. Zhang, F. Zhang, H. Zhang, Z. Zhao, J. Zhao, C. Zhao, Q. Zheng, Y. Zhi, Z. Zhou, B. Zou, K. Maier-Hein, P. Jäger, A. Kopp-Schneider and L. Maier-Hein

arXiv:2212.08568 2022.

DOI arXiv Cited by ~17

The number of international benchmarking competitions is steadily increasing in various fields of machine learning (ML) research and practice. So far, however, little is known about the common practice as well as bottlenecks faced by the community in tackling the research questions posed. To shed light on the status quo of algorithm development in the specific field of biomedical imaging analysis, we designed an international survey that was issued to all participants of challenges conducted in conjunction with the IEEE ISBI 2021 and MICCAI 2021 conferences (80 competitions in total). The survey covered participants' expertise and working environments, their chosen strategies, as well as algorithm characteristics. A median of 72% challenge participants took part in the survey. According to our results, knowledge exchange was the primary incentive (70%) for participation, while the reception of prize money played only a minor role (16%). While a median of 80 working hours was spent on method development, a large portion of participants stated that they did not have enough time for method development (32%). 25% perceived the infrastructure to be a bottleneck. Overall, 94% of all solutions were deep learning-based. Of these, 84% were based on standard architectures. 43% of the respondents reported that the data samples (e.g., images) were too large to be processed at once. This was most commonly addressed by patch-based training (69%), downsampling (37%), and solving 3D analysis tasks as a series of 2D tasks. K-fold cross-validation on the training set was performed by only 37% of the participants and only 50% of the participants performed ensembling based on multiple identical models (61%) or heterogeneous models (39%). 48% of the respondents applied postprocessing steps.